Files
orca/tests/e2e/cross-version-wire/structured-agent-session-surface-execution.ts
T
Brennan Benson 35e173a5b7 fix(native-chat): alert when a structured chat asks for approval or input mid-turn (#25766)
* fix(native-chat): alert when a structured chat asks for approval or input mid-turn

A structured Claude or Codex chat that stopped mid-turn to ask for approval or a
question raised no OS notification, no phone push and no unread dot: the host's
attention feed only fired when a turn settled. The host now announces each newly
pending prompt once, from the same feed, and owns the phone push for its own
structured sessions; desktops keep presentation only.

- host feed: a `prompt` edge per pending approval/question, opt-in on the
  existing stream; a clean settle that only restates an announced prompt is not
  re-announced to prompt-aware clients; failures always are
- host push: an in-process subscriber pushes the host's paired phones with
  labels from in-memory metadata, and withdraws a prompt's alert once answered
- one attention identity for desktop banner, phone push, dedupe and retirement;
  delivery dedupes on it instead of the per-workspace burst window
- agentSession.acknowledgeAttention routes reading a chat to its owning host,
  which retires what it pushed for that session via its dismissal record

* fix(native-chat): keep prompt alerts under their retirement identity

* Fix structured attention read and delivery lifecycle

* Allow later reads to retry failed attention retirement

* Prevent delayed prompt alerts after accepted reads

* Keep attention integration tests across process boundaries

* Check structured capability before cursorless attention reads

* Keep host retirement fixture in the Node typecheck boundary

* Keep attention test seams in the mock boundary

* Keep host notification formatting usable without Electron

* Declare agents in restored-host attention fixture

* Group session attention capability declarations

* Retire structured alerts on every read and when a remote prompt ends

- Mark read and Mark all read on a chat that was never opened, or is hidden,
  now cover what lit its row: the read boundary is the later of the transcript
  cursor and the newest edge the tab was handed, captured at the click. The
  one boundary closes the banners, retires relayed phone alerts and drives the
  host acknowledgement; an edge after the click stays live.
- An alert from a host older than journal cursors has no position, so a pane
  read retires it again, banner and relayed phone push, as before this change.
- A desktop-relayed prompt alert for a remote chat is withdrawn once that
  host's live status stops reporting attention (answered elsewhere, cancelled,
  host restart); cached status after loss of contact never settles it.
- The attention acknowledgement is agent-session.attention-ack.v1 with a
  required journal cursor and strict params; the cursorless no-op is gone.
- The dismissal store keeps a record whose origin a newer build wrote,
  dropping only that origin.
- Host reconciliation runs on restore and when a prompt leaves the pending
  set, not on every journal commit.
- Remove unused session-wide retirement (retire, retireMatching,
  retireMobileNotificationsMatching) and test-only store lookups.

* Settle relayed prompts only from host-dated status; reads default to view

- A remote chat's status and its prompt edges arrive on separate sockets, so a
  status row queued before the prompt could land after it and withdraw a
  relayed alert that was still pending. Settlement now needs a live,
  non-attention row whose host updatedAt (the journal's latest row time, never
  decreasing) is later than the prompt's host raisedAt. The owed prompt is
  cleared only once the settle call succeeds, so a failed call retries on the
  next qualifying row.
- An acknowledgement without captured reads now reads as a view. Mark read,
  Mark all read, the hover-card jump and both dashboard acks pass 'explicit';
  Activity's automatic per-turn read stays a view.

* Re-judge relayed prompt settlement after each call; dashboard watching is a view

- Settlement is one check of the current mirror, run on each status row, after
  each settle call returns, and when a prompt edge arrives. A row that lands
  while a call is out is no longer dropped, and an edge that arrives after its
  own resolution settles at once. A rejected call is retried only once the
  mirror holds a different row, so a failing call cannot loop.
- The Agent Dashboard's card click acknowledges as explicit; watching the open
  dialog acknowledges as a view, through both the in-window drawer and the
  pop-out relay, and no longer re-acks the state the click just read.

* Keep a reused read owner findable after its holders release it

A chat pane memoizes its read owner. When every holder let go in one commit and
the same owner was picked up again (React StrictMode's effect re-run in
development, or a pane remounted in one commit, as on a tab move), the release
dropped it from the registry and nothing put it back. The attention bridge then
found no owner, so a view read carried no cursor: the unread dot cleared but
the banner and phone alerts stayed. An owner now names its key again whenever
it gains an activation or a subscriber and the key is free, never over a
different live owner.

* Type the ack test's execution host ids

* Keep the turn-completion types in main's turn-completion wire module

Main moved these types into agent-session-turn-completion-wire.ts while this branch moved them into agent-session-attention-wire.ts. They now live at main's path with this branch's additions (journal cursor, prompt attention, prompt arm); the unused completion key is dropped.
2026-10-06 16:25:53 -07:00

113 lines
4.4 KiB
TypeScript

import { expect, vi } from 'vitest'
import { RuntimeSubscriptionRegistry } from '../../../src/main/runtime/runtime-subscription-registry'
import {
attachParams,
ATTENTION_READ,
paramsFor,
STRUCTURED_CALLS
} from './structured-agent-session-surface-manifest'
import type {
AgentSessionWireBuild,
RpcClientIdentity,
RpcReply
} from './versioned-agent-session-wire'
export function runtimeStub(overrides: Record<string, unknown> = {}): unknown {
const subscriptions = new RuntimeSubscriptionRegistry()
return {
getRuntimeId: () => 'runtime-1',
getClientSettings: () => ({ experimentalStructuredNativeChat: true }),
ensureStructuredAgentSessionHost: async () => undefined,
getStructuredAgentSessionCreateSupport: async () => ({ supported: true }),
structuredAgentSessionLaunchSeedOptions: () => undefined,
resolveStructuredAgentSessionCreateIntent: async () => {
const {
envelope: _envelope,
providerHandle: _providerHandle,
...resolved
} = attachParams(null)
return resolved
},
publishStructuredAgentSessionTab: () => {},
registerSubscriptionCleanup: subscriptions.register.bind(subscriptions),
registerOwnedSubscriptionCleanup: subscriptions.registerOwned.bind(subscriptions),
cleanupSubscription: subscriptions.cleanup.bind(subscriptions),
cleanupSubscriptionsByPrefix: subscriptions.cleanupByPrefix.bind(subscriptions),
...overrides
}
}
/** Every reply one call produced. Streaming methods answer more than once, and a
* refusal has to arrive as a reply rather than as silence. */
export async function callBuild(
build: AgentSessionWireBuild,
method: string,
params: unknown,
client: RpcClientIdentity,
runtime: unknown = runtimeStub()
): Promise<RpcReply[]> {
const replies: RpcReply[] = []
await build.createDispatcher(runtime).dispatchStreaming(
{ id: `request-${method}`, authToken: 'cross-version-token', method, params },
// oxlint-disable-next-line typescript/consistent-type-assertions -- SAFETY: the build's RPC dispatcher serializes its own reply; result/error shape is asserted by each surface case.
(raw) => replies.push(JSON.parse(raw) as RpcReply),
client
)
return replies
}
/**
* The one thing this suite exists to guarantee, written once and applied per
* build: every method the manifest declares is not merely registered but reaches
* its host method on this call, answers, and answers with its declared result.
*
* Written as a helper rather than inline because a build passing it is the claim,
* and each skew that registers the surface owes the same claim — a check that
* covers one method leaves the rest registered-but-unusable behind a green suite.
*/
export async function expectDeclaredSurfaceExecutes(
build: AgentSessionWireBuild,
hostCalls: Record<string, ReturnType<typeof vi.fn>>,
clientCapabilities: readonly string[]
): Promise<void> {
const retirement = vi.fn()
const runtime = runtimeStub({ retireStructuredAttention: retirement })
for (const { method, hostMethod, result } of STRUCTURED_CALLS) {
// Two methods share one host method, so "has been called" would already be
// true from the earlier one: only this call's own delta pins the pairing.
const before = hostMethod ? hostCalls[hostMethod].mock.calls.length : 0
const replies = await callBuild(
build,
method,
paramsFor(method),
{ clientKind: 'runtime', clientCapabilities },
runtime
)
if (hostMethod) {
expect(
hostCalls[hostMethod].mock.calls.length - before,
`${build.label}: ${method} did not reach the host`
).toBe(1)
}
if (method === 'agentSession.acknowledgeAttention') {
expect(retirement).toHaveBeenCalledExactlyOnceWith(ATTENTION_READ)
}
for (const reply of replies) {
expect(
reply,
`${build.label}: ${method} was refused: ${JSON.stringify(reply)}`
).toMatchObject({ ok: true })
}
if (result) {
// The declared answer, not merely a non-refusal: a handler that is
// registered and returns an execution error, or hands back someone else's
// envelope, fails here rather than passing as "reached the host".
expect(replies, `${build.label}: ${method} must answer exactly once`).toHaveLength(1)
expect(replies[0], `${build.label}: ${method} answered off-contract`).toMatchObject({
ok: true,
result
})
}
}
}