mirror of
https://github.com/stablyai/orca.git
synced 2026-10-07 08:02:21 +00:00
b407d06c1ecd9d237010984ac02d9cd3abe37df0
408
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
b407d06c1e | Reuse buffer cells during terminal cursor context scans (#25161) | ||
|
|
2b3b692d29 |
Reduce terminal test overhead while preserving full parity checks (#25151)
* Speed up terminal test oracles without reducing replay coverage * Call asynchronous parser through its checked test interface * Preserve evidence document final newline for concurrent merges |
||
|
|
ffd23fe25c | Avoid repeated runtime imports and recovery fixture seeding (#25155) | ||
|
|
d5bb69d348 |
Cut CI time in store, Git contention and readiness tests (#25147)
* Check store retention boundaries with a faster independent oracle * Replace CI diagnostic sleeps with gated contention and scoped transcript clocks |
||
|
|
136b990a05 |
fix(search): negotiate supported agents across mixed host versions (#25009)
Keep older request and reply parsers usable while current peers retain all supported history. Co-authored-by: nwparker <nwparker@users.noreply.github.com> |
||
|
|
af13da5db7 |
Recognize Build terminals and preserve Windows confirmation key routing
Recognize dsb terminal identity and preserve foreground Windows confirmation key routing. Co-authored-by: Wooseong Kim <innocarpe@gmail.com> |
||
|
|
2bd656c384 |
feat(native-chat): a Claude subagent waiting on a permission prompt reads as waiting (#22634)
* feat(native-chat): a Claude subagent waiting on a permission prompt reads as waiting
A subagent's permission request reaches the parent session's callback naming the
subagent that asked (agent_id) and the tool call it gates (tool_use_id). The
pending request is recorded with the asking agent. On every drain the child-work
producer re-derives which children a pending request blocks and hands that set
to the Claude child decoder, the one owner of each child's live edges: a blocked
child reads waiting on every live edge it reports, and a child that starts or
stops waiting is a live edge of its own. Answering, denying or cancelling the
request returns the child to its prior live state; nothing is stored beyond the
pending requests.
A live task_updated carrying an error now reaches the record as the child's last
message, without an ending or a new state.
The replay test drives a scrubbed capture of the real CLI (foreground allow,
deny, interrupt, background allow, and the main agent's own request) through the
real adapter into the host's child records.
* test(native-chat): a subagent's request names it before its tool call is read
* docs(agent-status): a subagent asking for approval waits in every lane; the parent row keeps the session's own attention
* test(native-chat): hand canUseTool the asking agent without widening the helper's cast
* test(native-chat): an interrupted Claude subagent settles cancelled, not failed
A captured interrupt shows the spawn call's error result ("The user doesn't want
to proceed…") arriving before the subagent's own `task_updated {status: killed}`.
The spawn result ends nothing (the child ends only on its own terminal frame), so
the child stays live until its `killed` status settles it cancelled. A genuine
failure, captured with the subagent on a model that does not exist, sends its
`failed` status before the error result and still ends failed. Both captures now
replay through the real adapter into the host's records.
* fix(native-chat): a Claude subagent's prompt makes the parent row wait, not block
A subagent's pending prompt made the whole session `attention`, which reads as
the main agent's own `blocked` and outranks the fold's waiting arm, so the
parent row read blocked where a CLI Claude parent reads waiting. The main
agent's state now reads only its own pending prompts.
- A Claude prompt row carries the linkage of the agent that raised it: the one
the permission request names, or the owner of the tool call it gates. The
same join decides which child reads waiting, so the two cannot disagree.
- The status summary projects the session's own status from root prompts only;
every other reader (delivery gates, teardown, restart) still asks whether
anyone is waiting on a human.
- An answer keeps the prompt row's linkage by the journal's own rule: a
revision that names no producer keeps the row's existing one.
- The child-tool queries gain the prompt's producer, so a prompt row and a
child record answer "which agent" from the same join.
* test(native-chat): say which ids the permission capture scrubs and which are its own
* fix(native-chat): the status clock dates attention by the session's own asks only
The session's status is now `attention` only for its own pending prompt, so the
clock's fallback to a subagent's ask could no longer be reached, and it read the
journal by a different rule than the status it dates. Both now read root prompts.
The journal also stamps a Codex subagent's prompt with its thread (#22532), so a
Codex child's approval is that child's wait in the Codex lane too. Two tests
written for the earlier rule are updated: a subagent's ask leaves a running
session `working` on its turn's clock, and a Codex child's answered approval
leaves the settled parent's Activity row done with nothing unread.
* fix(native-chat): a completion still says the user is asked when a subagent asks
The turn-completion feed marked a completion `awaitingUser` from the status
summary's `attention`. The status now means the session's own agent is waiting,
so a subagent's pending approval stopped reaching the completion. The projection
now also says whether anyone is waiting on the user, as the delivery gates,
teardown and restart ask it, and the completion reads that.
The waiting-subagent replay answers its prompt with the adapter's current
response shape.
* revert(native-chat): a live Claude task's error stays out of the child's last message
No capture shows a live task_updated carrying an error, and it is unrelated to
a subagent waiting on a permission request; it leaves this PR.
* test(native-chat): settle the Claude session's startup before replaying a subagent's request
A startup frame drained child work during the first await, so answering a
request freed the child even with the answer's own republish removed.
* fix(native-chat): a Claude subagent's prompt row names it as its other rows do
The prompt row stamped only the asking agent's id, so a nested subagent's
request lost the agent that spawned it, its spawn call and its run. It now takes
the linkage the asker's own rows take: the gated tool call's, when that names the
same agent, else the one resolved through the agent's spawn call. The provider's
agent id stays the asker's id.
* fix(native-chat): a subagent's request makes the parent row wait without a child record
The parent row learned that a subagent needed the user only from that
subagent's child record, so a request no record carried (a Codex child the host
never registered, a Claude task past the live cap) left the row working or done
while the approval card sat in the chat.
"Someone in this session must answer" is now one derived session fact. The
projection names two facts instead of a mode flag: the main agent's own status
(attention only for its own request) and structuredAgentSessionAwaitsUser (any
pending prompt). The status summary publishes the second as an optional
awaitsUser, and the shared fold reads it: the main agent's own ask is blocked,
otherwise awaitsUser or a waiting child record makes the row wait. Every caller
picks the fact it means: the completion edge's awaitingUser and the delivery
gate read awaitsUser; the quit snapshot folds the same two inputs the sidebar
does.
* fix(native-chat): a client that predates awaitsUser still reads a subagent's request as attention
A status summary's status is now the main agent's own, so a client built before
the split would read a subagent's request as working (or idle) and fold it with
code that has no awaitsUser input. Clients advertise
agent-session.status-awaits-user.v1; at agentSession.subscribeStatus the host
sends any client that does not the pre-split summary: attention whenever
awaitsUser is set, without the main agent's own tool line, verdict and clock.
The feed and every in-process reader keep the canonical summary. Transitional,
like the turn-item downgrade.
* test(native-chat): a Codex subagent's approval makes its settled parent's Activity row wait
The test pinned the parent row done while a Codex child asked, through a harness
that fed no child records, so it proved nothing about the ask. It now drives the
ask twice through the real host status store: with no child record (the
session's awaitsUser alone) and with the child's own record waiting from
thread/status/changed. Both read waiting with needsAttention while the ask is
open, then done with nothing unread.
* docs(agent-status): a subagent's request reaches the parent row through awaitsUser in every structured lane
The store reference said a Codex child's request still read as the main agent's
blocked and that only the Codex hook lane fed a waiting child. Both structured
lanes stamp the asking child and feed child records, and awaitsUser carries the
request when no record does. The liveness comment goes back to main's: a child's
blocked is a failed task on an older host's legacy rows.
* fix(native-chat): the restart dialog still headlines a subagent's pending approval
The quit snapshot now records the main agent's own state, so a subagent asking
while the main agent worked recorded `working` and the dialog said "Was
mid-reply" where it used to say "Waiting for your approval". The headline now
comes from the snapshot's pending prompt, whoever raised it, with the existing
copy; `state` stays the main agent's own.
* test(orchestration): a subagent's pending approval holds structured mail delivery
Scoping the delivery gate to the main agent's own request left every gate test
green; a subagent's request now has its own case.
* fix(native-chat): a subagent's request is dated by when it was raised, on every client
Since the summary's clock became the main agent's own, nothing dated a wait
that only a subagent's request held: a pre-split client was sent attention with
no clock, where the old host dated it by the subagent's prompt, and a new
client's waiting row fell back to the time it first saw it, so after a reload a
request the user had already read could read unread again.
The session fact is now when someone started being asked: awaitsUserSince, the
oldest pending prompt whoever raised it, and its presence is what awaitsUser
meant. A row waiting on someone else's request takes that as its clock; the
downgrade for a client without the capability dates its attention by it, which
is what the old host published. A cross-version test pinned to the last
pre-split release runs the same journals through that release's projection and
through this one plus the downgrade, and compares the whole summary. The Codex
end-to-end test also reads the host's own status row, and keeps a read ask read
through a later row and a reload.
* test(native-chat): the pre-split parity check compares only the fields the split owns
An additive summary field is safe for old clients, so comparing whole summaries
against the pinned release would redden on one. The wire comment now says how
the downgrade dates attention: the main agent's own oldest ask, else
awaitsUserSince.
* test(runtime): an aged host-held working summary states that nobody is asked
The test built its working summary by overriding the status of a published
approval summary, which still carried awaitsUserSince, so the row correctly
read waiting. It now drops the request as its scenario says.
* fix(native-chat): the chat's subagent block says waiting when the strip does
While a Claude subagent's request was open, the sidebar and the composer strip
read waiting but the subagent block in the chat history a few pixels above
still read "Kicked off 1 subagent working": it shows the journal's roster
state, and the journal records no wait.
The structured chat now hands its transcript the subagents the strip shows
waiting, read from the host's child records through the strip's own row model
and matched by the provider id the roster names each one by. A running entry
the host says is waiting reads waiting in the group row, its entry and its
section head, with the strip's word and the question colour; it reads the
journal's state again as soon as the host stops reporting the wait.
* fix(native-chat): a collapsed subagent group shows a wait beside a failed sibling
A failed sibling took the group row's one alert slot, so a group with a waiting,
a working and a failed child read "1 working +1 failed" and hid the wait; it
now reads "1 working +1 waiting +1 failed". The waiting set keeps its identity
while a child frame changes no wait, so the transcript's subagent rows do not
re-render on every frame, and the test of a wait ending now updates one mounted
row instead of remounting it.
* refactor(claude): one needs-input state on the parent; the asking subagent alone reads waiting
Drop the split of the main agent's own status from a session-wide "someone must
answer" fact: awaitsUserSince, the agent-session.status-awaits-user.v1
capability and its old-client downgrade, and every reader change that only
consumed them (fold, equality, ingest, delivery gate, turn-completion feed,
quit snapshot, resume headline, status clock, status bridge, attention
dispatch) go back to main. The parent row again reads one needs-input state
for a pending request whoever asked, dated as before.
Kept: a request's owner recorded once on its prompt row with full producer
linkage; the asking subagent's own record reads waiting, re-derived on every
update; the chat history's subagent block reads that same state; an answered
subagent request stays in its subagent's group.
A subagent now waits only on a request the user can still answer (its card
open, no answer underway), and the adapter frees it before the host records an
answer or dismissal. So a waiting child record always sits beside the pending
card, and main's fold never reads the parent as waiting on it: no window after
an answer, and no ~3 s wait after a card dismissed by Stop.
* fix(claude): a subagent waits only beside its committed card
A subagent's wait was pushed to the host as soon as its request arrived,
while the request's card row reached the journal at least a microtask later.
So every subagent request published the parent row as waiting before
blocked (the main agent's own fold reads a waiting child that way), and
Activity got an extra unread "waiting" event that main never shows.
The card is now the one record of an open request. The translator records
the asker on the card once (its row's linkage) and counts the card open only
after the sink confirms its rows landed, then publishes the wait; anything
that closes the card (an answer underway, a dismissal handed to the host,
Claude's own withdrawal, the session's end) frees the subagent first. So
every publish that shows a subagent waiting also shows its pending card, and
the parent reads one needs-input state, exactly as on main.
This retires the registry's view of pending requests (unclaimed(), the
asking-child join) and the translator's holdsOpen. The prompt row's linkage
takes one rule: the agent the provider names, else the gated call's owner.
The parent-row proof now runs through the real deferred sink, durable
journal and status feed, publishing as production does, and checks at every
publish that waiting subagents have pending cards and that the parent row
matches a host fed no waits.
* fix(native-chat): a closed sink's dropped writes never read as landed
The sink's written() resolved ok when the sink was closed with writes still
queued, so a subagent's prompt card could count as open with no row in the
journal. written() now reports a close that dropped writes admitted so far
as not landed; drained() and lifecycleBarrier() keep reading a closed sink as
settled.
Tests: a card never opens when its sink closes first; a card Claude withdraws
while the sink holds the cancelled row back closes at once; two subagents
asking at once, and the main agent asking beside a subagent, keep the parent
row as before with each waiting subagent beside its own card; a process that
dies mid-request leaves no subagent waiting.
* fix(claude): a withdrawn subagent request frees its child before its card closes
Main's sink now hands each write to the journal as it is submitted, and an idle journal commits it
and runs the publication at once. Claude's own withdrawal of a subagent's request therefore closed
the card and published the parent row before the child's wait was freed, so one publish showed the
subagent waiting beside no pending card (fg-interrupt replay). The child's wait now also requires the
request to still be open in the registry, and a withdrawal republishes child work before the
journal takes the close.
* refactor(claude): trim subagent request waiting to the common pattern and its essential tests
The chat history no longer marks a subagent block as waiting: the approval card itself carries the
request, and the asking subagent's row in the sidebar and composer strip reads waiting, as before.
NativeChatWaitingSubagentsProvider, native-chat-waiting-subagents.ts and their renderer changes go.
A subagent waits while its request is still open and unanswered in the prompt registry and its card
has landed in the journal. The registry check also covers a withdrawal under backpressure, so the
card list no longer filters pending cancellations itself.
Tests: one integration file replays the captured CLI frames through the real adapter, sink, journal
and status feed (renamed claude-subagent-permission-request.test.ts), with the asking subagent's
state timeline, attribution, nested linkage and a card write that waits for the journal. The
producer-harness waiting test, the redundant prompt-card cases, the harness reducer swap and three
unused captures (deny, interrupt, main agent, failed subagent) are removed.
* test(claude): pin the parent row's dating when a subagent asked first, and narrow the oracle's claim
|
||
|
+34 |
c2dba12072 |
Add host-owned OpenCode and Devin managed account profiles (#24636)
* Add host-owned OpenCode and Devin account profiles * Manage OpenCode and Devin profiles in account Settings * Expose registered account roots to host transcript readers * Clarify managed profile provider flags * Retain isolated Electron home in browser sidecars * Restore inherited account environment and preserve cleanup retries * Check relay environment values before merging * Consolidate managed account type imports * fix(accounts): use existing localized provider names Align the new Japanese account copy with the existing catalog repair policy. * Keep managed account baselines private to the execution host * Align account enrollment help with accepted providers and flags * Show the active System account in managed profile lists * Document the validated Linux managed account scope * Fix managed account removal, credential audits, and runtime bundling * Quarantine removed accounts and audit captured credential snapshots * Continue Antigravity IDE and 2.0 history in new CLI conversations (#24692) * feat(antigravity): bridge IDE history into new CLI conversations * fix(antigravity): preserve fresh-launch model and environment for IDE references * fix(antigravity): forward IDE history opt-in through desktop IPC * fix(antigravity): rebuild remote IDE reference startup on its host * fix(antigravity): register IDE continuation action labels * fix(antigravity): confine IDE references and bound metadata reads * fix(antigravity): localize IDE continuation badges * Preserve scanner service cache assertions and refresh Antigravity opening metadata * Preserve Antigravity opening joins and target folder runtime authority * fix(opencode): retry timed-out SSH plugin updates (#24666) Preserve bounded retry behavior and the current-main status-envelope fields. Original-PR: #24124 Reviewed-source: |
||
|
|
aca2d51e0e |
fix(jcode): harden Windows hooks and negotiate remote history (#24998)
Redirect the managed Windows payload file into curl instead of starting pipeline shells, register native Windows delivery coverage, and document Jcode v0.89.0+ as the upstream launcher requirement for invisible hooks. Negotiate Jcode history in both directions with mixed-version Orca hosts, preserving supported search filters and old-client response compatibility. Co-authored-by: czzczz <chanzrz_zbf@foxmail.com> Co-authored-by: JianJia2018 <39438074+JianJia2018@users.noreply.github.com> |
||
|
|
4c2cb3cf12 |
fix(jcode): quote the managed hook path, drop tools from patch prompts
Three fixes, all on paths this PR could not exercise locally. The managed hook command was stored as a bare path. jcode tokenizes that string shell-style before exec'ing it directly (parse_hook_command, crates/jcode-terminal-launch/src/lib.rs): unquoted whitespace splits, and every unquoted backslash is consumed as an escape. So on Windows `C:\Users\me\.orca\agent-hooks\jcode-hook.cmd` reached exec as `C:Usersme.orcaagent-hooksjcode-hook.cmd` and no hook fired at all, and a POSIX home with a space split into two arguments. Store the path single-quoted (verbatim, backslashes included), falling back to double quotes for a path containing a single quote. Existing bare entries are already repointed by the stale-key path, and getStatus accepts both forms so the repair is not reported as a user-owned hook. The quoting helper was previously dead code that only tests called; the three production sites now use it. isJcodeManagedCommand also normalizes separators, since a `/`-only needle never matched a Windows entry. Commit-message generation feeds a staged patch to `jcode run` as the prompt — attacker-influenced text — while jcode's default profile exposes shell, read, write, and MCP. Pass `--tool-profile none`, which resolves to an empty allowed-tool set in jcode's config (base_allowed_tools), matching the read-only posture claude (plan) and codex (read-only) already take. docs/reference/jcode-hook-events.md was never actually in this PR: the repo ignores docs/** and tracks reference docs by allow-list only, so the captured-payload evidence four source comments point at was silently dropped. Allow-list it. Co-authored-by: czzczz <chanzrz_zbf@foxmail.com> |
||
|
|
668d6d45c4 | Reuse measured Electron preparation for current Terminal Perf refs (#24968) | ||
|
|
ec47f0c41b | Drain removal fixture jobs before resetting and deleting their records (#24977) | ||
|
|
97fd7edb39 |
fix(editor): recognize Ruby task and configuration files
Add Ruby task/configuration filenames to the existing generated language associations. Co-authored-by: ggbdpq <ggbdpq@gmail.com> Co-authored-by: Neil <neil@stably.ai> |
||
|
|
f93808dfff |
Collect test-selection evidence when full unit tests fail (#24955)
* Collect advisory unit-selection evidence from failed full runs * Trigger checks after retargeting the evidence fix to main |
||
|
|
30e0ccaba6 | Let retired-cache GC observations settle across the existing six-turn budget (#24967) | ||
|
|
8f26bfad22 |
Use the measured pnpm lookup policy automatically in hosted root CI (#24951)
* Select lookup automatically for the measured hosted root-install profile * Record hosted automatic-mode cold cache publication proof |
||
|
|
6e6f651380 |
Avoid repeated pnpm archive downloads in cache producers (#24927)
* Let measured cache producers keep stores without downloading them * Check that restore-only callers do not publish a producer path * Enable the measured producer mode and record hosted comparisons |
||
|
|
9e27050955 |
Add verified native Antigravity Accounts on the owning runtime (#24691)
* Add verified native Antigravity accounts on the owning runtime * Keep Antigravity usage tied to its observed native account * Refuse oversized encrypted Antigravity snapshots before writing * fix(antigravity): localize account heading and search terms |
||
|
|
1d5d02e396 |
Resolve explicitly configured command aliases for workers (#24648)
* feat(orchestration): resolve explicitly configured command aliases * docs(orchestration): explain configured command aliases * fix(orchestration): validate configured aliases with target shell grammar * fix(orchestration): use actual shell and refuse assignment-only aliases * test: preserve typed calls in configured worker target checks Replace Reflect.apply with the existing typed prototype call pattern so the unchanged regression cases pass the anti-slop lint gate. |
||
|
|
75f8f34ce3 | Overlap independent Linux headless runtime builds (#24910) | ||
|
|
1241ce1b49 |
Skip slower root package-store restores in macOS PR jobs (#24908)
* Skip slower root package-store restores in macOS PR jobs * Update the companion cache-policy contract |
||
|
|
533446dde6 |
Stop mocked renderer imports from qualifying headless CI (#24902)
* Decouple headless running-work tests from the renderer * Keep the shared running-work probe contract documented |
||
|
|
a824fb74ab |
Reuse the headless detector compiler without installing full dependencies (#24895)
* Reuse the headless detector compiler without full dependency setup * Keep optional compiler-cache saves from failing cache warming |
||
|
|
58cf72d48e | Skip slower root package-store restores in Linux PR jobs (#24896) | ||
|
|
add1c55590 |
Skip slower Windows root package-store restores in CI (#24885)
* Skip slower Windows root package-store restores in CI * Update reviewed mobile dependency-store cache expression |
||
|
|
61836f6026 |
Reduce scheduled CI cache warming to every six hours (#24881)
* Reduce scheduled CI cache warming to every six hours * Document cache warmer recovery interval and measured tradeoff |
||
|
|
cc73c8e1a7 | ci: overlap ARM SSH setup and independent observation waits (#24714) | ||
|
|
ac28e8c85e |
Skip dependency installation for known headless build inputs (#24716)
* ci: defer headless dependency installation until graph analysis is needed * docs: align headless CI rollout with platform and cache policy * test: isolate headless detector output from the parent CI step |
||
|
|
9503f9de34 |
fix(runtime): settle Codex tui-idle waits on its hook done, leaving working rows to the rules (#24541)
A headless orca serve has no window to write the Codex ready title, so a Codex tui-idle wait settled only after three quiet seconds. Rule files gain profile.hooks: "turn-end": a fresh hook done settles the wait, while a working or permission row leaves the decision to the rules, so a Codex whose Esc posts no event (before its Interrupt hook) cannot hang the wait. Codex moves from identity-only to turn-end. |
||
|
|
cd8d03bc06 | fix(dsh): recognize 0.2 profiles and open workspace composer (#24589) | ||
|
|
13ecf051c3 |
Reuse prepared Windows native builds in SSH CI (#24555)
* ci: reuse qualified Windows server slots for SSH host tests * ci: reuse prepared relay addons after an exact native cache hit |
||
|
|
efbf651c7b |
Reduce CI setup costs and fixture failures (#24537)
* Let scheduled CI warmers wait and measure WebRTC startup * Measure a smaller daemon shutdown fixture image * Counterbalance WebRTC startup and verify retained fixture files * Record CI fixture measurements and remove temporary pilots * Clarify fixture build dependency cleanup evidence * Make coalesced snapshot fixture delivery deterministic * test: type the PTY write delay observer |
||
|
|
6153fbcfe4 |
Reduce redundant headless server CI work (#24527)
* ci: avoid unrelated headless server qualification * ci: skip headless detection for ineligible draft PRs * ci: preserve cross-host qualification and skip supplied prerequisites * ci: include Windows server cache validation in change detection |
||
|
|
5f308bfa9c |
revert: take the 26 Phase 3 (#16741 port) PRs back out of main (#24559)
* Revert "feat(orcad): source-side dormant export of a relay-hosted SSH target (#16741 T6-8) (#24519)" This reverts commit |
||
|
|
34ae0933e4 |
fix(worktrees): keep creation fast in large repositories (#24346)
* fix(worktrees): remove repeated scans and keep prepared checkouts fresh * fix(worktrees): reclaim unlocked fallback preparations safely * refactor(worktrees): simplify creation ownership and idle maintenance * fix(git): keep ref maintenance armed after an index-only pass An idle attempt that found the pack index due but refs still cooling down returned without rescheduling, so loose refs from the arming fetch waited for the next write instead of the ref cooldown. |
||
|
|
cdfdadf9ea |
fix(runtime): settle tui-idle on hook state for agents whose hooks cover the whole turn (#24388)
* fix(codex): install Codex's Interrupt hook so an Esc-cancelled turn settles
Codex 0.150+ fires an Interrupt hook when the user presses Esc on an
approval prompt or mid-tool, and nothing else. Orca did not install it, so
the pane stayed blocked/working until the next prompt.
- Add Interrupt to the managed Codex events and label maps, written with
Codex's 3s cap (a larger value triggers a startup clamp warning).
- Hash the timeout Codex hashes (Interrupt is clamped to [1,3], default 1)
so self-computed trust matches Codex; pinned against a real 0.159.3 hash.
- Map a root Interrupt to the existing cancelled-turn record
(markCodexLeadTurnInterrupted), keeping child work in the fold; a
child-scoped Interrupt is ignored. Relayed rows take the same path.
* test(runtime): add a readiness census pinning every tui-idle verdict
Replays every recorded agent PTY transcript frame by frame through a real
runtime pane (agent-known and agent-unknown, clocked and clockless) and a
synthetic evidence matrix for all 43 TuiAgents, and compares each verdict
and tui-idle wait outcome to committed run-length-encoded baselines.
Refs STA-9098
* test(runtime): pin the census quiet probes to literal windows
A census that read TUI_IDLE_QUIESCENCE_MS would move with it; fixed 2999/3000 ms
reads and a fixed 2000 ms poll step make a changed window show as changed verdicts.
Refs STA-9098
* test(runtime): say which census probe writes runtime state
Refs STA-9098
* refactor(codex): let the hook builder own Codex's per-event timeout
The managed hook's timeout is now Codex's own normalization of the shared
budget, and every installer derives its trust entry from the hook it wrote,
so no installer repeats the Interrupt special case.
Claude-Session: codex-interrupt-hook review
* refactor(codex): route Interrupt through the Stop lead update with an outcome
Interrupt now writes the lead record through the same setCodexMainAgentTurnState
call as Stop, so markCodexLeadTurnInterrupted keeps its original signature.
Drops the child-scoped Interrupt guard: Codex never runs Interrupt hooks for
subagents and its input schema has no agent_id.
Claude-Session: codex-interrupt-hook review
* test(runtime): observe the census through settled panes and caller-visible waits
- Read each verdict through the runtime's own settle seam (evaluateTuiIdleForLeaf) instead
of re-wiring evaluateTuiIdle/leafTuiIdleEvidence/buildTerminalWaitText, so the census is
coupled to one runtime method, not to the module STA-9098 rewrites.
- Let the runtime finish each chunk (one macrotask turn) before reading. The old read raced
work chained on the paint, so 14 frames pinned a microtask-ordering artefact.
- Record when a wait settles (@start vs @poll), not just its outcome.
- Exit each pane's PTY after reading it so its emulator is freed.
- Replace the hand-grouped families, literal fixture list and per-pane split flag with a
directory-scanned catalog, one baseline per replayed pane, and size-balanced shards.
- Run the synthetic matrix in one file; it takes about 2 s.
* test(runtime): cross dialog-versus-ready-screen order with every title in the census matrix
Blocked detection is position-ordered (design doc 11.5): the later of a blocker and a ready
anchor wins. The matrix now paints a workspace-trust dialog after, and before, each agent's
ready screen under every title, so a rule engine that loses that ordering fails per agent.
* test(runtime): read the census baseline field without Reflect.get
The anti-slop lint rejects Reflect.get on parsed input.
* refactor(runtime): read Antigravity, Cline, Prime Agent and Cursor readiness from rule files
Adds agent-state-rules/: a zod-validated JSON file per agent, one priority list of
screen rules per agent (idle with strength and requiresQuiet, or hold), and text
anchors that feed the shared, position-ordered blocked layer every pane reads first.
The three screen-ruled agents and Cursor's approval menu and prompt move to data;
the Antigravity text scan stays code as a named anchor. Their old code paths are
deleted. Every other agent still runs through the existing lanes, unchanged.
The readiness census baselines are untouched and pass.
Refs STA-9098
* test(runtime): cover the agent state rule engine's schema, priority, rows, anchors and lanes
Refs STA-9098
* fix(runtime): refuse rule patterns that repeat an optional or alternating group
The load-time regex check only flagged a repeated group whose body held * + or {,
so (a?)* and (a|aa)+ passed though both backtrack exponentially. A repeated
group's body must now be fixed: no quantifier of any kind and no alternation.
The comment states the remaining polynomial gap instead of claiming linearity.
* refactor(runtime): give agent state rules and text anchors one when/answer shape
Every rule and text anchor is now when (a region and what it must show) plus
answer, each a discriminated union, so part (b) adds title, text and status
regions and working or blocked answers as new variants instead of new fields.
- Cursor's prompt is two anchors answering working and idle; the one-off
workingIfAfter and followedBy fields become a general after test.
- Anchor literals and the probe banner must be lowercase, since they are
matched against the lowercased tail.
- screenProbeBanner moves under profile, the place for non-detection facts.
- why is required on every rule and anchor.
- A blocked anchor must name a lastOf literal, which the prefilter keys on.
* docs: point the readiness evidence docs at the agent state rule files
* refactor(runtime): read Codex, Claude, OpenCode, Pi, OMP and Gemini readiness from rule files
The rule engine gains the regions and answers these agents need, as closed-list entries:
- rule regions `title` (the classified title status) and `text` (one of the file's idle text
anchors, settled), and a `predicate` form of the screen region for named engine scans;
- `withoutClock: skip` for strong quiet rules a clockless pane must not believe;
- anchors (renamed from textAnchors) gain a `title` region, and `live` and `hold` answers;
- `profile.screenSource` (trusted grid or live screen), and an `unknown-pane` file for panes
with no known agent.
Codex's header, composer and provisional-startup checks become named predicates referenced
from codex.json; its ready header, header and startup hold become shared text anchors. Native
idle title markers become shared title anchors; name-only title handling becomes each agent's
idle-title rule. The agent-specific branches in terminal-wait-detection.ts and
tui-idle-evidence.ts are deleted, and the "later live prompt cancels a blocker" rule now reads
only rule-file anchors (plus Muse, which moves in part b2).
No behaviour change: the readiness census baselines are untouched and pass.
Refs STA-9098
* test(runtime): cover the rule engine's title, text and predicate regions and the bundled anchors
Refs STA-9098
* fix(runtime): reject a rule file that repeats an anchor or rule id
A text rule names its anchor by id, so a repeated id let a file pass validation and then throw
while compiling. Also states that engineVersion bumps once a version ships; version 1 is still
being defined.
* refactor(runtime): fold the working anchor answer into live
The engine treated an anchor's working and live answers identically: both mark a live prompt
that cancels an earlier blocker and settles nothing. Cursor's busy prompt now answers live, so
anchors have one non-settling prompt answer.
Refs STA-9098
* refactor(runtime): read the shared π title anchor from pi.json alone
Pi and OMP paint the same `π - <session>` rest title, and title anchors apply to every pane,
so one copy covers both.
Refs STA-9098
* refactor(runtime): key every rule file and read the trusted screen from screenSource alone
readsTrustedScreen no longer also asks for a screen rule (every trusted file has one, and the
schema requires screenSource where it matters), so rule-less files need no filter. A rule's
match is a plain boolean, and compileTitleAnchors is module-private.
Refs STA-9098
* test(runtime): pin that a clocked Codex pane takes no other agent's ready text
No test failed when holdsReadyTextToQuiet was removed; this one does.
Refs STA-9098
* fix(agent-hooks): keep an OMP approval wait until omp resolves it
omp posts tool_execution_start a few milliseconds after
tool_approval_requested, while its Approve/Deny select still holds the
human. Both mapped onto the pane row, so the working event overwrote the
blocked one and the pane read as busy for the whole prompt.
A working event now leaves an OMP approval wait in place; only
tool_approval_resolved or a new turn ends it. An ask row is unchanged:
its own tool_execution_end ends it. The test replays the order a live
omp 17 run posted for a denied bash call.
Refs STA-9100
* refactor(runtime): select the fresh hook row on any of a terminal's handles or pane keys
selectFreshExplicitAgentStatus matched one handle and one pane key and
returned only the mapped status. The row selection now takes sets of
handles and pane keys, an optional received-at floor, and returns the
row itself, so a reader can see the main agent's own state. The old
function keeps its signature and result on top of it.
Refs STA-9100
* feat(runtime): let tui-idle read hook state for agents whose hooks cover the whole turn
tui-idle read no hook state. Hook state reached readiness only through
the `<Agent> ready` titles the window writes, so a headless `orca serve`
never saw it (#16095), and Codex settled only once its screen had been
quiet for three seconds.
Rule files gain `profile.hooks: "authoritative" | "identity-only"`,
defaulting to identity-only. Codex (with its Interrupt hook), OpenCode,
OpenCode 2, Pi and OMP are authoritative. For them a fresh hook-store
row decides ahead of every other lane:
- the main agent's turn decides (`mainAgent.state` when published), so a
subagent's Stop does not end the lead turn: done settles strong,
working holds, a permission wait never settles;
- the tail's blocked text goes through the existing permission arbiter
with the turn as its explicit status, so a denied prompt's dialog left
in the tail no longer blocks a turn the hook says ended;
- the row joins on every pane key and terminal handle the PTY owns.
No row, a stale, restored or other agent's row, a session-start done,
and a row from before a PTY respawn all fall back to today's lanes. That
keeps startup on the screen and text rules: Codex posts SessionStart
only with the first prompt. Claude, Cursor, Gemini and the rest stay
identity-only.
The readiness census has no hook server, so its frames are unchanged.
Refs STA-9100
* docs(agent-status): record readiness as a reader of the hook store
Refs STA-9100
* fix(runtime): ignore a hook done older than the latest input Orca wrote
A finished turn leaves a fresh `done` row. A caller that sends the next
prompt and waits at once could settle on it before the new turn's first
hook arrives, so the wait returned while the agent was starting work.
Orca's own input writes (terminal send, agent prompts, mailbox pointers)
now stamp a per-PTY input clock, and the hook lane reads no `done`
received before it; the pane falls back to the screen and text rules
until the agent reports again. A `working` row is unaffected.
Refs STA-9100
* docs(agent-status): note the input floor on the hook lane's done
Refs STA-9100
* fix(runtime): take the hook lane's input floor from the PTY run's input record
The hook lane ignored a done older than Orca's latest write to the pane, kept in a
new per-PTY map stamped by a wrapper threaded through four write sites. The PTY
run register already sits on both write funnels, so it now records the last
input (launch writes included, terminal replies not) and the lane reads it.
Keys the user types now count too, which closes the restart-in-the-same-shell
gap: typing `codex` to relaunch no longer lets the previous process's done read
ready while the new one boots.
The respawn floor moves from the shared row join into the lane, beside the
input floor; the freshest row predates a floor exactly when every row does.
* test(runtime): drop runtime hook-lane cases the unit suite already proves
Working over a ready title, a permission wait, and an identity-only agent are
decided inside evaluateTuiIdle and covered there; the runtime suite keeps the
wiring: the join, both floors, Pi's own OSC 133 markers and the arbiter.
* fix(runtime): record a PTY's last input even when main adopted it without a spawn commit
A materialized pane re-adopted by the renderer returns before the spawn-commit
site, so it had no run record and its input never moved the hook lane's floor.
The last input now lives beside the run records: any PTY's input counts, and a
new process's commit still clears it.
* fix(runtime): keep a running process's input time when main reattaches or adopts it
A reattach or adoption commit without an incarnation id cleared the PTY's
last-input time, so a prompt sent just before an SSH adoption was forgotten
and the hook lane could accept the previous turn's done as ready. Only a new
process (or a reattach naming a different incarnation) now starts clean; the
first-input fact follows the same rule.
* docs(runtime): say why a lead turn that ended reads ready while a subagent runs
* fix(runtime): refuse uppercase contains terms in text anchors, which read the lowercased tail
A text anchor's after and lines tests run on the lowercased tail, so an
uppercase contains term loaded and then never matched. Build the text test
schema from the literal it accepts and give anchors the lowercase one. Also
drop a probe-banner early return that no bundled catalog reaches.
* refactor(runtime): state Codex's provisional startup and title anchors as plain rules
The provisional-startup hold becomes a lastOf anchor with an all/none test, so
its TypeScript scan goes. Title anchors drop their status field (every caller
already gates on an idle title), and withoutClock keeps only the value a rule
can set.
* fix(runtime): leave Codex readiness to its title and screen rules
Codex before its Interrupt hook posts nothing for an Esc mid-turn, so its hook
row stays working and a hook-authoritative tui-idle wait hangs until the row
goes stale. Current Codex already settles fast through its ready title.
Co-Authored-By: Claude <noreply@anthropic.com>
---------
Co-authored-by: Claude <noreply@anthropic.com>
|
||
|
|
8ff6296bc7 |
Speed up serializer checks and keep native caches stable (#24476)
* Reuse serializer oracle cells and isolate native cache policy * Preserve native cache post-save paths and record hosted oracle gain * Record native cache reuse and separate cancel-test startup budget |
||
|
|
e3cb32791e |
refactor(runtime): read four agents' readiness from JSON rule files through one engine (#24348)
* test(runtime): add a readiness census pinning every tui-idle verdict
Replays every recorded agent PTY transcript frame by frame through a real
runtime pane (agent-known and agent-unknown, clocked and clockless) and a
synthetic evidence matrix for all 43 TuiAgents, and compares each verdict
and tui-idle wait outcome to committed run-length-encoded baselines.
Refs STA-9098
* test(runtime): pin the census quiet probes to literal windows
A census that read TUI_IDLE_QUIESCENCE_MS would move with it; fixed 2999/3000 ms
reads and a fixed 2000 ms poll step make a changed window show as changed verdicts.
Refs STA-9098
* test(runtime): say which census probe writes runtime state
Refs STA-9098
* test(runtime): observe the census through settled panes and caller-visible waits
- Read each verdict through the runtime's own settle seam (evaluateTuiIdleForLeaf) instead
of re-wiring evaluateTuiIdle/leafTuiIdleEvidence/buildTerminalWaitText, so the census is
coupled to one runtime method, not to the module STA-9098 rewrites.
- Let the runtime finish each chunk (one macrotask turn) before reading. The old read raced
work chained on the paint, so 14 frames pinned a microtask-ordering artefact.
- Record when a wait settles (@start vs @poll), not just its outcome.
- Exit each pane's PTY after reading it so its emulator is freed.
- Replace the hand-grouped families, literal fixture list and per-pane split flag with a
directory-scanned catalog, one baseline per replayed pane, and size-balanced shards.
- Run the synthetic matrix in one file; it takes about 2 s.
* test(runtime): cross dialog-versus-ready-screen order with every title in the census matrix
Blocked detection is position-ordered (design doc 11.5): the later of a blocker and a ready
anchor wins. The matrix now paints a workspace-trust dialog after, and before, each agent's
ready screen under every title, so a rule engine that loses that ordering fails per agent.
* test(runtime): read the census baseline field without Reflect.get
The anti-slop lint rejects Reflect.get on parsed input.
* refactor(runtime): read Antigravity, Cline, Prime Agent and Cursor readiness from rule files
Adds agent-state-rules/: a zod-validated JSON file per agent, one priority list of
screen rules per agent (idle with strength and requiresQuiet, or hold), and text
anchors that feed the shared, position-ordered blocked layer every pane reads first.
The three screen-ruled agents and Cursor's approval menu and prompt move to data;
the Antigravity text scan stays code as a named anchor. Their old code paths are
deleted. Every other agent still runs through the existing lanes, unchanged.
The readiness census baselines are untouched and pass.
Refs STA-9098
* test(runtime): cover the agent state rule engine's schema, priority, rows, anchors and lanes
Refs STA-9098
* fix(runtime): refuse rule patterns that repeat an optional or alternating group
The load-time regex check only flagged a repeated group whose body held * + or {,
so (a?)* and (a|aa)+ passed though both backtrack exponentially. A repeated
group's body must now be fixed: no quantifier of any kind and no alternation.
The comment states the remaining polynomial gap instead of claiming linearity.
* refactor(runtime): give agent state rules and text anchors one when/answer shape
Every rule and text anchor is now when (a region and what it must show) plus
answer, each a discriminated union, so part (b) adds title, text and status
regions and working or blocked answers as new variants instead of new fields.
- Cursor's prompt is two anchors answering working and idle; the one-off
workingIfAfter and followedBy fields become a general after test.
- Anchor literals and the probe banner must be lowercase, since they are
matched against the lowercased tail.
- screenProbeBanner moves under profile, the place for non-detection facts.
- why is required on every rule and anchor.
- A blocked anchor must name a lastOf literal, which the prefilter keys on.
* docs: point the readiness evidence docs at the agent state rule files
* fix(runtime): refuse uppercase contains terms in text anchors, which read the lowercased tail
A text anchor's after and lines tests run on the lowercased tail, so an
uppercase contains term loaded and then never matched. Build the text test
schema from the literal it accepts and give anchors the lowercase one. Also
drop a probe-banner early return that no bundled catalog reaches.
|
||
|
|
f69052e113 | Reuse qualified Windows server builds and dependency verification records (#24448) | ||
|
|
43d9b43d3f |
feat(ssh): remote orcad stop by request file and journaled decommission (#16741 T6-4) (#24449)
Clients stop an orcad that advertises health.stopRequests through its slot-local request file and keep SIGTERM for older builds. Decommission runs through the activation journal and fence: it refuses while the terminal census is live or uncounted, stops the instance with an instance-bound managed request, cancels a stop orcad never acted on, and deactivates the record only on proven exit. orcad gains --cancel-managed-stop and an exclusive per-transaction decision file so a cancel can never race a dispatched stop. POSIX-only and inert: no production caller. Co-authored-by: m4air <m4air@m4airs-Air.localdomain> |
||
|
|
b093d3ab20 |
feat(orcad): supervisable server: stop requests, managed stop receipts and a lifetime that keeps its lock on failed teardown (#16741 T6-3) (#24433)
orcad stops through slot-local and instance-bound request files, so a reused PID is never signalled. A managed stop is proven by its completion command and recorded as a receipt. Optional daemon retirement is best effort: an idle daemon retires, while a busy or unverifiable one stays up with its admission fence released. Runtime teardown runs in reverse order and keeps the instance lock and profile admission when any writer fails to stop. Browser discovery no longer delays readiness. Legacy worker recovery and watcher children are drained before the final flush. Headless terminal close no longer waits on a renderer tab that does not exist. No production deployment. Co-authored-by: m4air <m4air@m4airs-Air.localdomain> |
||
|
|
197ea3a3b3 |
Free PR CI capacity by avoiding repeated setup and real-time test waits (#24355)
* Reduce repeated PR setup and transcript timing waits; add hosted comparisons * Align parallelism contract with Node-only external rebuild toolchain * Record hosted coverage and launch package, store, and cancellation comparisons * Apply hosted Windows setup savings and remove measured test waits * Keep measured PR package gains and remove completed comparison jobs * Report measured test counts with precise units |
||
|
|
c6cfcc034e |
refactor(native-chat): structured chat failures always reach the diagnostics log (#24312)
* refactor(native-chat): give the structured chat host one required logger The structured chat runtime took an optional onError callback that the desktop never passed, so a late dispatch settlement, an unanswered-dispatch release, a journal event-sink write and a provider lifecycle delivery that failed were dropped with no trace. Other host failures went to scattered console.warn calls, which reach nothing in a packaged desktop build. The runtime and host now take one required logger (warn/error with a scope and fields). The production logger writes each entry as a failed span to <userData>/logs/main.trace.ndjson, which the diagnostic bundle collects, and to the console (stderr under a supervised headless host). The runtime and the host wrap it so a logger that throws never fails what it reports, and the install refuses without one. Sites that deliberately kept a recovery-capsule error out of the log still log no error object. * refactor(native-chat): hand the chat host's collaborators the logger, and give orcad its trace file The delivery loop, idle sweep, queued-message drain, lease renewer, event sink, conversation map and provider start/exit settlement each took an internal error callback that the host mapped onto the logger. They now take the logger itself and log under their own scope. The event sink keeps one onFailed hook, which decides whether to stop the provider, not whether to report. The dead-generation settlement returns its failure so each caller logs it under its own scope. orcad now installs the desktop's local trace sink under its own data root, so a headless host's chat failures reach <data-root>/logs/main.trace.ndjson as well as stderr. Also passes the logger in the test fixtures the first commit missed, which tc:node caught. * fix(native-chat): keep repeated chat failures from flooding the trace file, and record their causes - The production structured-chat logger writes a repeated failure (same level, scope, session, message and error text) once per 5 minutes, carrying how many repeats it swallowed; the tracked set is capped at 256. - Trace entries now carry the error's code (and SQLite errcode) and up to three causes by name and message. - A chat read whose conversation will not open is logged through the host's logger (open-for-read), and so are the runtime's chat-tab bookkeeping failures that already hold the host. - orcad writes its own orcad.trace.ndjson, closes it after every quit handler, and flushes it on process exit; a trace file that cannot be opened leaves tracing off instead of stopping the app or orcad. - Tests: the desktop wiring test proves the logger reaches the trace sink, and the privacy tests read every level the logger received. * fix(native-chat): log a created chat's tab-publication and launch-prompt failures through the host's logger * fix(native-chat): key a repeated chat failure on everything its entry writes The repeat suppression keyed on the message and the error's text, so two refusals with the same code but different causes, a plain error and a refusal of one code, or two object-valued errors shared a key and the second was swallowed for five minutes. The key is now the entry's whole written content (fields, code, errcode, refusal reason, cause chain, a stable rendering of a non-error value) plus the error's name and message; a refusal's reason is also written. * test(native-chat): pin that an error's name keeps two repeated failures apart * test(native-chat): build the refusal in the repeat-key test as the wire does * fix(native-chat): read an error's code and a refusal's reason by narrowing, not Reflect.get |
||
|
|
6d1a97ef98 |
fix(ssh): launch the Windows relay outside sshd's job so standard users work (#24224)
* fix(ssh): launch the Windows relay outside sshd's job without WMI Win32-OpenSSH kills a session's job on close but allows breakaway. relay.js gains a one-shot launcher mode that starts the detached relay with CREATE_BREAKAWAY_FROM_JOB through the staged process-tree addon, so a standard user no longer needs a WMI Remote Enable grant. WMI stays as the fallback for a relay without the addon, and a refusal there is named. The Windows SSH-host lanes drop their WMI grant and assert the breakaway route and adoption. * fix(ssh): find runtime holds without WMI on a standard-user Windows host The store GC read held runtimes through Get-CimInstance Win32_Process, which WMI refuses to a standard user's SSH logon, so the pass kept every runtime. On a refusal it now reads this account's own process image paths through Get-Process. * build(relay): ship the Windows relay launcher addon in every desktop package macOS and Linux packages carried Windows relays without windows-process-tree.node, so a legacy-runtime relay they uploaded to a Windows SSH host could not launch outside sshd's job and fell back to WMI, which a standard user is refused. A reusable Windows job now compiles the x64 and arm64 addons once and uploads them; release-cut, release-mac-build, and the hourly/daily/adhoc mac builds download them before build:release and require both arches. Staging now rejects a binary with the wrong PE machine, the ReadProcessMemory import, or no spawnOutsideJob export, so a stale pre-launcher build cannot ship. * ci(ssh): run the Windows SSH-host lanes when the relay process-tree build scripts change The staging and gyp-rebuild scripts decide which windows-process-tree addon the relay ships, so a change to either must re-prove the Windows host cells. * test(ci): find the mac orcad-template download by artifact name The release mac job now also downloads the relay Windows process-tree addons, so the first download-artifact step is no longer the template's. --------- Co-authored-by: m4air <m4air@m4airs-Air.localdomain> |
||
|
|
14d4bb2e2a |
fix(ssh): Windows hosts without Add-Type staging; runtime-store GC on Windows (#24149)
* fix(ssh): collect the pinned-Node runtime store on Windows hosts
Windows SSH hosts now run runtime-store GC instead of skipping it: one
PowerShell inventory reads .runtime-ref-node-<sha> and .runtime-node refs from
every version dir, and one Get-CimInstance Win32_Process query filtered on an
image path under runtimes\ adds process holds (never by image name; a failed
query keeps everything). Stale upload stages are swept with the same rule as
POSIX. Promotion and the post-upload hold check now take the store lock on
Windows too, and the lock's own commands run unwrapped there.
Windows relay version-dir liveness now honours .relay-pid (design D5): a live
PID answers ALIVE before any pipe is touched, a dead one (ESRCH) plus refusing
pipes is exited, anything else is unverifiable. The runtime probe adopts a
pinned node.exe an earlier vault reader left without a .verified marker after
running it.
* fix(ssh): Windows stage fencing and vault runtime go through the verified node.exe
Upload-stage file identity on Windows no longer compiles an Add-Type P/Invoke
helper when the relay runs on Orca's verified pinned node.exe: the stage
commands run a fixed fs.lstatSync(..., {bigint:true}) script through it. It
prints the legacy helper's vol:high:low lowercase hex, and identity files are
compared after normalising hex spelling, so old and new clients recover each
other's stages. Host-Node relays keep the legacy helper; the choice is
documented in windows-edr-posture.md.
The Windows OpenCode vault reader now installs the pinned runtime through
ensureRemoteOrcadNodeRuntime (official zip, host-side extraction, .verified,
store lock) instead of uploading a client-extracted node.exe, and the relay dir
gains a .runtime-ref-node-<sha> so store GC keeps the runtime the vault uses.
* test(ssh): run the Windows stage-identity and store-GC tests on the Windows lane
The legacy/node.exe identity compatibility test and the Win32_Process hold path
were gated to win32 but no CI lane ran them. Add both files to the Windows
package lane and a real running-node.exe hold test.
* test(ssh): tear down Windows-lane temp trees through removeTreeSync
* test(ssh): grant the store lock to the Windows OpenCode runtime setup test
The Windows promote now runs under runtimes/.store-lock, so the mocked host
must answer the lock's CreateNew step.
---------
Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
Co-authored-by: m4air <m4air@Mac.localdomain>
|
||
|
|
bd90da7a5b | ci: share PR planning setup and reuse the static native cache (#24329) | ||
|
|
ddd4927a0b |
build(orcad): server node-pty slots at glibc 2.28, plus a glibc 2.17 compat slot (#24134)
* build(orcad): build server glibc slots on glibc 2.28 and add the glibc 2.17 compat slot
Design D6: the default linux-{x64,arm64}-glibc node-pty slots now build in
manylinux_2_28 (digest-pinned) and are gated at glibc 2.28 / GLIBCXX_3.4.25
through a floor profile on verify-linux-glibc-floor.cjs; the desktop keeps
its Ubuntu 20.04 (2.31) default.
Adds the opt-in linux-x64-glibc217 compat target: NODE_RUNTIME_COMPAT_ASSETS
pins the unofficial glibc-217 Node (update/check pin scripts cover it, outside
SERVER_TARGETS), and a new CI lane builds the compat slot in manylinux2014
with static libstdc++, gates it at glibc 2.17 with no shared C++ runtime in
DT_NEEDED, and smokes it under the glibc-217 Node.
* refactor(node-runtime-pin): route compat lookups through isCompatServerTarget; keep the glibc doc's slot-name paragraph intact
---------
Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
|
||
|
|
6593d7d194 |
feat(orcad): run orcad on the pinned Node instead of Bun (#24110)
* ci(daemon): gate PRs on daemon protocol crossing from the newest release Lands daemon-protocol-facts.mjs from the Windows update diagnostic branch with a stricter parser, and adds check-daemon-protocol-crossing.mjs (rule R1): the working tree must attach the newest release tag's daemon. Rollback crossing is reported only. Runs in the cross-version-wire job, which already has full tags; tag selection moves to config/scripts/stable-release-tags.mjs so both use one rule. * feat(persistence): run profile backups in the worker whenever its entry is bundled * refactor(orcad): make profile and native preflight runtime-neutral The profile preflight parser now takes the expected runtime identity from the caller (shipped callers pass the pinned Bun identity), and the native preflight is renamed to orcad-runtime-native-preflight with neutral wording. * feat(runtime): pin the Node 24.21.0 server runtime with an offline CI check Add src/shared/node-runtime-pin.ts (NODE_RUNTIME_PIN, SERVER_TARGETS, NODE_RUNTIME_ASSETS for all 8 server targets plus the headers tarball), generated by config/scripts/update-node-runtime-pin.mjs from the nodejs.org and unofficial-builds SHASUMS. check-node-runtime-pin.mjs verifies, with no network, that the pin tracks the locked Electron, matches engines.node's major, and covers exactly SERVER_TARGETS; it runs in the static analysis job. ORCAD_BUN_TARGETS consumers now read SERVER_TARGETS so there is one target list; orcad's Bun runtime and build output are unchanged. * test(persistence): skip plain-Node backup selection tests in the Bun profile suite * fix(runtime): reject a pinned archive that belongs to another target * ci(daemon): fail PRs that swap a runtime launcher and bump the daemon protocol D7.1 R3: hosting orcad or the daemon on another runtime is not a protocol change, so one PR must not do both. The launcher file list lives in the check script; the allow-runtime-launcher-protocol-bump label overrides it. * feat(orcad): select pinned-Node slots by a .runtime-node marker D7.1 R5: a Node slot names its shared runtimes/node-<sha256>/node through .runtime-node instead of .build-target, so Bun-era clients read it as a legacy slot rather than exiting 78 on a missing bun-runtime. Nothing builds the marker yet. * fix(runtime): load the Node pin without the typeless-module warning check-node-runtime-pin.mjs now requires the pin and takes nodeDistArchiveName from its own module, so it no longer loads the update script's build graph. * fix(orcad): resolve Node slots to the design's runtimes/node-<sha>/bin/node layout * feat(orcad): 8-slot node-pty prebuilds against the pinned Node headers at N-API 8 - build-orcad-prebuilds.mjs adds win32-x64/arm64 (conpty.node, the vendored conpty.dll/OpenConsole.exe, upstream's N-API conpty_console_list.node), compiles in a scratch copy against the hash-verified pinned headers (node.lib pinned per Windows arch) with NAPI_VERSION=8, rejects post-8 node_api_* imports, and writes a schema 2 manifest with per-file sha256, N-API level and the glibc need. - --require-slots [slots] verifies files against hashes; --smoke loads the slot under the pinned Node and spawns a PTY; --print-slot names the host slot. - The slot installer gates on N-API, libc, arch, glibc and file hashes instead of the exact NODE_MODULE_VERSION, and installs nested files (conpty/). - bun-profile-tests.yml builds, verifies and smokes each runner's slot. * fix(orcad): scope node-pty's glibc .symver pins to glibc on musl prebuild slots musl's unversioned libc cannot satisfy openpty@GLIBC_* references at link time, so the Alpine slot compile would fail. Pin the staged pty.cc guard to __GLIBC__ and assert both musl transforms against the installed patch. * feat(orcad): run orcad on the pinned Node instead of Bun A packaged orcad slot now references the pinned Node 24.21.0 by its executableSha256 (`.runtime-node`, `.server-target`) instead of carrying bun-runtime, and ships node-pty from the slot's prebuild, only its own ripgrep, and no Windows Bun PTY gate. The runtime lives beside the slots at runtimes/node-<sha>/bin/node (node.exe on Windows, upstream name). - build:orcad (build-orcad-node.mjs) builds the host slot's prebuild when missing and places the pinned runtime; the template is schema 3 with per-target files. - handoffToBundledOrcad() resolves the slot's runtime reference and checks process.versions.node against the pin; a host Node >= 18 still hands off. Startup preflight keys on running as that runtime; callers expect 'node'. - orcad and its daemon use node-pty (ConPTY + windows-pty-job on Windows); the Bun PTY sources, gate entry and canUseBunPty branches are removed. - SSH deploy uploads the official archive once per pin, extracts and hash-checks it on the host, and self-tests it before publishing. Bun slots stay launchable for rollback; Node slots never use host Node. - The runtime materializer is generic over pinned assets; the Bun wrapper remains only for the OpenCode vault reader (design Phase 2). - Cross-runtime test: a profile DB written by Bun 1.4.2 (WAL left by SIGKILL) opens and backs up under the pinned Node, and the reverse. No daemon PROTOCOL_VERSION change (design D7.1 R3). * docs(ci): name the headless lanes after the pinned Node, drop Bun shard timings Design D10: ci-demand-rollout.md and ci-runner-efficiency.md follow the bun-profile-tests.yml -> node-server-tests.yml rename; shard timings drop the deleted Bun PTY tests and follow the renamed ones. * chore(ci): count the runtime archive download as a runtime launcher path * fix(orcad): pin the macOS C++ standard for node-pty prebuilds The official Node headers' config.gypi sets clang: 0, so common.gypi skips its gnu++20 xcode_settings and Apple clang 15 (macos-14 runners) compiles node-addon-api as C++98. * fix(orcad): resolve the preflight's slot through realpath, as the handoff does A symlinked orcad.js handed off to its real slot's pinned Node, but the startup and profile preflights read the symlink's directory, found no runtime marker there, and silently skipped the readiness check. * refactor(ssh): drop materializeCachedNodeRuntime, which nothing calls Deploys upload the verified official archive (design D5); no client path needs an extracted Node executable cached by digest. * test(orcad): gate the Bun-to-Node upgrade and Node-to-Bun rollback with live terminals Design D7.1 R1/R3/R4 and D7.2. The last Bun orcad and this checkout's Node slot are installed side by side under ~/.orca-remote, launched and stopped with the client's own deploy commands, and share one data root. Each direction proves the incoming orcad adopts the outgoing runtime's daemon (same PID, same shell, output continues), opens its profile database and backs it up with its own shipped worker, and that GC keeps the slot the live daemon was forked from. The node-server Linux lanes provide Bun 1.4.2 and build that Bun orcad from main, and run with --cross-runtime. --artifact and --cross-runtime now make their tests fail on a missing input instead of skipping. * ci(node-server): pin node:24.21.0-alpine by its multi-arch index digest * test(ssh): name the runtime archive fixture after its role * test(node-server): load node-pty from the packaged slot in artifact runs The node-server lane installs dependencies without building node-pty, and Linux has no upstream prebuild, so the real-PTY failed-I/O teardown test (picked up by the pty-subprocess selector) could not load pty.node. In --artifact runs, alias node-pty to out/orcad's shipped slot so the test exercises the addon orcad actually runs under the pinned Node. * fix(orcad): let the Windows profile preflight exit after its PTY probe On Windows, node-pty keeps the conout worker thread and pseudoconsole alive until kill(), even after the shell exits. The PTY health probe never killed a cleanly exited probe, so the packaged preflight printed its readiness line and then hung until the build's 30s timeout, reported with an empty stderr. - The probe kills its PTY on Windows after exit and uses the bundled ConPTY the daemon spawns with. - The preflight exits once stdout is flushed; its owner reads to EOF. - Preflight failures now report code, signal, timeout, stdout and stderr. * test(node-server): load the slot's node-pty in the real-PTY test, not by alias A vite alias redirected only ESM imports of node-pty; windows-pty-job and local-pty-utils resolve it through require, so Windows loaded two conpty.node copies and the Git Bash job-membership proof read an empty job. The failed-I/O teardown test now loads node-pty through a fixture that picks the packaged slot in artifact lanes. The pty-subprocess selector was a prefix that also pulled in its POSIX-host sibling unit tests, which pr.yml runs and which were never qualified on Windows. Select the directory plus the two sibling files that belong here. --------- Co-authored-by: m4air <m4air@m4airs-Air.localdomain> |
||
|
|
daf63e659c |
fix(runtime): read Antigravity, Cline and Prime Agent readiness from the live screen (#24222)
* fix(runtime): decide Antigravity readiness from the live screen agy paints its composer with cursor addressing, so the line-folded wait text misses the 1.2.14 accept-edits and plan composers and an ended turn, while the grid keeps the bare `>` caret painted mid-turn and behind the /model picker. Read the screen's bottom rows instead: rule, caret, rule, `? for shortcuts`. A clocked pane is held to quiescence (tier 1b) because the submit repaint reads ready for a moment; a clockless restored pane settles from the screen alone. When a trustworthy screen exists it decides, so the name-only title lane no longer settles an open picker. Retires the visible-read probe's Antigravity branch: the probe now runs the shared screen rule for any screen-ruled agent without an output clock, and keeps its generic empty-pane read for everyone else. Adds twelve agy 1.2.14 recordings and a replay suite shared by screen-ruled agents. STA-8741. * fix(runtime): decide Cline readiness from the live screen Cline paints its composer box with cursor addressing on the alternate screen, so no text rule saw it and worker-start timed out at agent_readiness (#23268). Read the box off the grid: rule, an empty composer with one of the captured placeholders, rule, the Plan/Act row and the auto-approve row, with no braille spinner above it. A streaming reply repaints the same box once its spinner has scrolled away, so Cline is tier 1b only: a clocked pane waits for quiet and a clockless one never settles from the screen. The screen now decides for a Cline pane, which shuts the quiet-process lane that would have settled its unworded tool-approval prompt and the Cline Desktop promo. readLiveTerminalScreenLines now returns raw rows: the read projection blanks a composer it takes for a draft, and it takes Cline's placeholder for one, so a typed draft and an empty composer looked the same. Adds nine cline 3.0.66 recordings (macOS) and the 3.0.65 Windows capture from #23269. STA-8741. * fix(runtime): decide Prime Agent readiness from the live screen Prime redraws its composer on the alternate screen, so the text tail never showed a settled prompt and tui-idle timed out (#22153). Read the grid instead: a bare `>` directly over the `<- manage` footer, with no braille status row (`Writing - 6s`) above it. The footer and caret alone stay painted for a whole turn. Replayed chunk by chunk, Prime erases that status row before redrawing it, and on first launch paints the idle composer just before the trace-sharing question covers it. Both keep repainting, so a clocked pane is held to quiescence (tier 1b); a clockless restored pane settles from the screen alone. Adds nine prime-agent 0.9.8 recordings (isolated HOME, OpenRouter) and the two 0.9.5 captures from #22154. STA-8741. * refactor(runtime): drop Cline-only readiness branches Cline now follows the same pattern as Antigravity and Prime: a screen rule plus table entries. - Drop MID_TURN_COMPOSER_AGENTS. onPtyData stamps lastOutputAt on every chunk, so a re-attached streaming pane has an output clock from its first byte; the exception only guarded a pane that printed nothing since attach. A clockless Cline pane now settles from its screen like the other two. - Drop the 'ready-body' rest-signal entries for all three agents. The rest signal is read only by quietForegroundLane, and a readable screen already shuts that lane and the title lane (isReadinessDecidedByScreen), so the entries only removed the quiet-process fallback for a pane with no trustworthy grid. The census now checks that screen-shut instead. - Drop the Cline rule's auto-approve row check; no recorded verdict depends on it. Kept: raw rows from readLiveTerminalScreenLines. Every frame of every codex-* and qoder-* capture at 120x40, 80x24 and 100x32 gives the same isKnownReadyPromptBody (with and without a clock) and isQuietReadyScreenBody verdict through both readers. Serializer known-failures for the new captures are pre-existing serializer behaviour, not this branch: row-0 cells restore with a true-colour background where the source has the default (the DSH class), and Prime's cursor restores at column 119 instead of the pending wrap at 120 (the qoder class). STA-8741. * fix(runtime): trust a screen rule only on the PTY's own grid Review findings on the screen-ruled readiness (STA-8741): - A grid out of step with the PTY garbles cursor-addressed chrome, and a model resize does not make the TUI repaint. readLiveTerminalScreenLines now returns null unless the emulator's grid matches the PTY's reported size and was never reflowed without a repaint (a re-attach that learned the real size late), so the pre-existing lanes decide there instead of timing out. - The visible-read probe reads the draft-blanking projection, which turns Cline's `❯ Ask anything...` into a bare `❯`. It now restores the blanked composer row before the rule reads it; `terminal read --screen` output is unchanged. - The quiet lane no longer ORs the text rules over a trustworthy screen that refused; without one, tier 1 already ran them. No recorded verdict changes. Tests: ready recordings on a mismatched and on a reflowed grid settle through the old lanes; the restored-pane probe runs every ready recording through the real projection; the rest-signal census checks the lane verdict with and without a screen. * test(runtime): trim STA-8741 recordings to the screens they prove * refactor(runtime): one screen verdict for every screen-ruled lane readScreenRuledReady, readScreenRuledQuietReady and isReadinessDecidedByScreen each re-derived the same thing: the agent's rule applied to a trustworthy live screen. They collapse into readScreenRuledVerdict (true / false / null), which tier 1, tier 1b and the lane gate read. This also makes a refusal final in tier 1: a clockless pane whose trustworthy screen refused fell through to the text rules, so retained ready text could settle over an open picker (Greptile review). The quiet tier already refused there; now both do. The tier-1b agent set derives the screen-ruled agents from the rule table instead of listing them again, and the lane test that repeated the census case is dropped. * refactor(runtime): let the visible-read probe read its own output clock The probe's clock was captured at start and threaded through the wait dependencies as a one-off parameter. The probe now reads it from the live record when its screen read returns, which is also the fresher answer. * fix(runtime): trust a reflowed grid again once a PTY resize repaints it The reattach-reflow flag was never cleared, so a pane stayed on the old lanes for the rest of its life even after a real resize made the TUI repaint (Greptile review). The record now keeps the reflowed grid, and a PTY resize off that grid clears it; an echo of the same size sends no SIGWINCH and keeps it. Tests: the reflow case in every screen-ruled suite now includes a same-size echo, and an Antigravity recording only the screen reads ready settles after a resize and repaint. * refactor(runtime): keep screen-rule trust and raw rows to screen-ruled agents Two shared changes reached agents this PR does not target: the live screen reader returned raw rows, and it refused a grid that did not match the PTY. Both now live in readScreenRuledLines, which only the screen-ruled agents read (screenReader picks it from the rule table); readLiveTerminalScreenLines is main's again. The probe keeps main's Antigravity-banner trigger, so a Codex or unknown pane is probed exactly as before. Proof: the non-screen-ruled suites give identical pass sets on this branch and its base (1,781 tests), and replaying every other recording frame by frame through the readiness and blocked verdicts, for its agent and for an unknown pane, gives identical results (93 pairs). A new test keeps a Codex pane reading its screen when the PTY reports another grid; it fails if the trust check moves back into the shared reader. |
||
|
|
0b79720c2e |
feat(native-chat): the chat strip and the sidebar read the host's child records (#22614)
* feat(native-chat): publish the host's child records to the status summary and the chat strip
The status summary and the background-task channel now read a session's child
records from the host's canonical store, through the status sink its row
landed in, and derive the legacy task and subagent shapes from the same views.
The parent row folds its child-work liveness from those records at ingest,
not from the summary's task list. The adapters no longer push their task DTO
to clients: the onBackgroundTasksChanged path is gone, and a child-work ingest
is what republishes both the summary and the strip. Finished children stay
listed until the session's own next turn starts. A reader that predates child
views never receives a roster whose rows are all settled.
* feat(native-chat): the chat strip reads the host's child records with its parent's verdict
The strip's roster now renders from the child views its channel carries, and
passes the verdict the session's own status row gives its children, built the
way the sidebar builds it (the row's freshness and the status feed's
observation). So one child reads the same in the strip and the sidebar, live,
after the transport drops, and once the row goes stale. A roster of finished
children stays shown until the next turn but no longer animates the monitoring
indicator or blocks conversation commands.
Tests: an end-to-end run on a host with no renderer (a real hook server as the
status sink) shows the summary and the strip channel carrying the same records
at every step, the parent row folded from them, retention, and an older
client's task list holding live work only; a wired renderer test shows both
surfaces agree when live, lost and stale.
* test(native-chat): a newer host's view degrades, an older reader keeps its live roster, a finished roster holds nothing open
- The view decoder ignores unknown keys, degrades unknown kinds, states,
outcomes and memberships, and drops only rows it cannot identify.
- At the RPC boundary a reader that predates child views gets no strip for a
roster of finished children and never the views themselves; a stop-only
reader keeps rows whose host offers no targeted stop.
- The strip shows finished children without reading them as live work.
- The row keeps its child list's identity when a summary repeats it.
* refactor(native-chat): the summary's task list is the legacy projection's live rows, unfiltered
* test(native-chat): type the switch tests' mocks instead of asserting them
* test: remote clients advertise reading child views
* docs(agent-status): the structured row folds the store's child records
* fix(agent-status): keep the view reader's header from reading as a value import to the renderer boundary
The renderer node-builtin boundary test scans raw text, so a header comment
that said "imports" ahead of the import block turned the type-only import of
agent-status-child-work into a value edge that reaches node:crypto.
* refactor(native-chat): the status summary's broadcast equality gets its own module
The status feed crossed the file-size limit once the summary gained the main agent's turn
outcome beside the child views. Which summary changes reach every session list now lives in
structured-agent-session-status-summary-equality.ts.
* fix(native-chat): command admission reads the strip's child records
A conversation command was refused on the provider tracker's own roster
while the strip read the host's child records, so a drift between the two
rule sets could refuse /clear with a stop instruction the strip had no
button for. Admission now reads the same records through the same read as
the strip, uses the strip's liveness fold, and asks for a stop only when
the strip renders one. The adapter contract no longer exposes the tracker
roster, so no host decision can read it.
Also records when the legacy child shapes die, every earlier death of a
settled child, and the display-precision invariant behind the summary's
clock tolerance.
* refactor(native-chat): command admission takes only what it reads of a turn
* fix(native-chat): the session list drops a session's children when the store does
A session's end no longer removes its child records: a child still running
settles with an outcome nobody reported, and a finished one stays listed.
Records now leave only at the session's own next turn, at the cap on settled
records, or when the host lets go of the session and its row leaves the store.
The summary kept after the host lets go used to strip its children on close,
a rule of its own. It now re-reads them from the store when the row leaves,
through the same read every live summary uses, so the session list and the
chat strip list the same children at each step, including a forget with no
close. Closing only revokes ownership, as before the child records existed.
* test(native-chat): write the Codex frame script's parent row out step by step
Once every surface reads the child records, the provider tracker's roster is
no oracle: it and the records read the same child executions, so agreeing
with it cannot catch a defect in either. Each frame now states the child
liveness and the parent row it must fold to.
* fix(native-chat): the idle sweep and the restart snapshot read the host's child records
The idle sweep (keep an agent running while its subagents or commands run) and
the restart-resume snapshot (what a chat was doing when Orca stopped it) both
read the provider tracker's roster through the adapter interface, which no
longer carries it. Both now take the host's one child-record read, the same one
the status summary, the chat strip and command admission use.
The snapshot's working test also folded that roster through the shared fold's
old `backgroundTasks` input, which the fold no longer reads, so a settled lead
whose subagent was still running would have been offered nothing. It now hands
the fold the records.
* test(native-chat): the child-record tests follow the merged command lifecycle
A command is a live child record from its start and is removed, not settled,
when it stops, whatever Codex tagged it. The end-to-end switch now shows the
child's `npm test` as a live row beside its dev server, and both are gone once
they exit; only the finished subagent stays listed until the next turn. Letting
go of the session is its tab closing, since a closed conversation whose tab
remains keeps its row.
Command admission's finished row is a subagent, the one kind that settles, and
the failed-verdict row test admits its live subagent as a host record, the only
thing the row folds.
* refactor(native-chat): the status feed's journal projection cache gets its own module
The status feed crossed the file-size limit once the child records joined the
agent-start signal and the completion feed's status read. The per-journal
projection, cached per commit, now lives in
structured-agent-session-status-journal-projection.ts.
* test(native-chat): the admission test's compaction resolves with a real outcome
Main's compaction result is a tagged outcome; the host-level admission test
resolved its mock compaction with an empty object.
* feat(native-chat): the sidebar lists running subagents; the strip, running then the newest finished
The host keeps every child record; what each surface lists is picked from them on every read, so
nothing is stored twice. The status summary, which every session list reads, now carries only
running children (and a finished one whose shell still runs, which reads monitoring): a finished
or failed subagent leaves the sidebar and stays in the chat's strip. The strip lists every running
child, then the newest finished ones, 100 rows in all; more than 100 running all show.
This matches common practice: sidebars show live subagents, and finished ones stay in the chat's
panel, newest first. No wire field is added. An older client reads fewer rows: its legacy task
lists were already live-only in the summary, and the strip's settled tasks come from the same
bounded roster.
* fix(sidebar): one rule for what the worktree sidebar lists: running children, from every source
`worktreeSidebarListsChild` is the one definition: a child that runs, counting a finished one
whose own shell still runs (it reads monitoring). The sidebar's row builder applies it to every
child source it reads, a terminal agent's hook roster and a chat session's records alike, and the
host's status summary applies the same predicate, so the sidebar's payload stays small. The chat's
strip keeps finished children, newest first.
A terminal agent's hook roster already drops a child on its own stop, so nothing changes there:
a teammate between turns and a child gone quiet still run, and still show. The selection module
moves to `agent-child-work-listing.ts`, since it now covers every source, not only chat sessions.
* test(native-chat): the switch test passes the startup child key main's status bar takes
* fix(native-chat): a finished child stays until the user's next send, not a turn Claude opens on its own
Claude wakes the agent on its own when a background task ends, and that wake is a
new root turn. Keying retention on the newest root turn retired every finished
child about two seconds after a background agent or shell finished, so its
outcome never showed in the strip.
Retention now keys on the user's newest send the provider accepted (a message, a
steer or a command), with the journal epoch so a rewind still retires. A turn the
provider opens itself and a subagent's turn carry no send. Replayed captured wake
orders through the real adapter, hook server and status feed.
* fix(native-chat): a background Stop reaches the tasks the child records show
The strip draws a row's Stop, and /clear, /compact and rewind wait for background
work, from the host's child records, but the Claude adapter still resolved which
tasks a Stop reached from its own tracker's roster, and refused to stop at all
once that roster was empty. A task the records kept live after the roster dropped
it showed a Stop that sent nothing and blocked those commands until the chat tab
closed.
The host now resolves the provider ids a Stop sends from the records (the same
per-row rule the strip and admission use; every such row for stop-all), and the
adapter stops exactly those, with no tracker guard. An acknowledged stop ends the
record: a running task sends its own stopped frame first, and the CLI answers
success with no frame for a task it no longer knows. A refused stop leaves the
record live. No production code reads the tracker's roster any more.
* fix(native-chat): one rule for a finished child that still owns live work, at any depth
The listing kept a finished child whose work ran through any depth of ownership,
but retention at the user's next turn protected only the direct owner, so a
finished agent whose finished subagent still ran a shell was removed and that
subagent jumped to the top level. Both now read settledOwnersOfLiveWork.
* fix(native-chat): an older client sees a Codex child's shell as it did before views
Clients that predate child views read a flat task roster derived from the views.
It listed a Codex child agent's shell as an extra row beside the running agent,
then as a bare command once the agent finished. The derivation now hides a
running agent's commands and names a finished agent's as "<agent> — <command>",
as the Codex tracker did; the label rule moves to a shared module both use.
* fix(native-chat): the chat decodes a roster's child rows once, as the frame arrives
The client reducer compared raw wire rows, so a row shaped by a newer host could
throw there, and the strip decoded a new array on every render, which defeated its
grouped-rows memo while a turn streamed. Rows are now decoded where the frame
enters the reducer, an unchanged roster keeps its identity, and the strip's parent
context is rebuilt only when one of its values changes.
* fix(native-chat): the strip channel forgets a closed conversation's roster
It kept the last roster fingerprint of every conversation for the host's lifetime.
The conversation map now tells observers when one leaves it. Also corrects the
summary's children comment: it carries running children only.
* docs(native-chat): rewrap the retention comment
* fix(native-chat): a task's own ending replaces a Stop's, and a child finished after the user wrote stays
Two lifecycle gaps from the round-1 fixes.
A Stop acknowledged ahead of the task's own ending relabelled it. The SDK hands
Orca a control answer as soon as it reads it and queues other frames, so when a
task finished just as the user pressed Stop, the acknowledgement arrived before
the task's completion the CLI wrote first, and the task read "Stopped" with its
result lost. An acknowledged Stop now ends a record provisionally (outcome basis
`stop-acknowledged`); the task's own terminal frame replaces it, and nothing
replaces an ending the task reported itself.
A finished child still vanished with no new action from the user when the send
the provider took was written before the child finished: a steer Claude takes at
its next boundary, or a queued draft handed over at the end of the turn. The
user's next turn now carries when they acted (the send's written time, or the
draft's queued time, both on the host clock), and only children that finished at
or before it retire; a later one stays until the user's next send. A rewind
still retires every finished child.
Also: a Stop-all keeps stopping the remaining tasks after one request fails, then
reports the failure.
* fix(native-chat): the strip keeps one empty list for a roster that omits one
A roster with only running or only finished rows made a new empty array on every
render, so the strip regrouped its rows each time. One shared empty list keeps
its memo.
* fix(native-chat): a strip row whose owner the 100-row budget cut renders under the main agent
The budget can keep a finished child and cut its finished owner. The child still
named that owner, so it rendered nowhere. The selection now clears an owner it
did not keep, as the view contract says for an owner outside the projection.
* fix(native-chat): a stop-all that times out stops asking, and "no children" is sent once
A stop-all kept asking after a request timed out, so a Claude CLI that stopped
answering control requests cost one full deadline per task, while the chat's
sends, Stop and /compact waited behind it. A timeout now ends the loop; other
request failures still let the remaining tasks be stopped.
The chat strip channel never remembered that it had sent "no children", so
every change in a chat with none re-sent that frame to each subscriber. It now
remembers it, and forgets only when the conversation closes.
Also renames agent-child-work-stop.ts to agent-child-work-stop-targets.ts, which
says what it answers: the provider ids a background Stop reaches.
* fix(native-chat): the status feed reads a journal snapshot with no submissions, and e2e tests use the child-work reader
CI on
|