Files
orca/config/scripts/verify-mobile-web-app-bundle.mjs
T
Brennan Benson 761d63a4e5 feat(agent-launch): keep long prompts off the launch line and paste them after readiness (step 1 of 7) (#24257)
* feat(agent-launch): host-side prompt delivery for agent.launch

The host's agent.launch typed any launch prompt into the shell as part of the
launch command. A long or multi-line prompt then ran line by line in the
shell, and an agent that never showed readiness or crashed at startup had
nothing guarding where its text went.

agent.launch now carries a prompt on the typed line only when the line stays
one line, control-free and at most 512 bytes; otherwise the agent starts
clean and the host pastes the prompt once the agent's own ready signal fires
(bracketed paste plus its composer marker or a quiet render, read only after
the shell's last hand-off, never while the pane's own shell is proven in
front), with main's draft-paste bytes and an Enter 50 ms later. Orchestration
worker starts wait on tui-idle as before. A replay-safe launch admits and
claims its ledger row in one write, Qwen Code gets a second Enter, the
desktop and phone share one launch-refusal classifier, and hosts advertise
agent.launch.prompt-carry.v1.

Split out of #23748, which moves the desktop source-control buttons onto
this path.

* fix(agent-launch): keep a short-lined multi-line prompt on a local zsh launch line, as main did

#24257 moved every multi-line or over-512-byte prompt off the typed launch line and pasted it after readiness. The phone's AI buttons and review notes, whose multi-line prompts main typed whole into zsh, then reached Claude 0.5-3 s later and their RPC reply waited for the paste.

The host now names the shell a local macOS or Linux line is typed into, the way the spawn picks it, and a multi-line prompt rides a zsh line when every line is at most 512 bytes and the whole line at most 8 KB. A real-zsh test types such a line through Orca's own ready barrier and startup write, including when a slow user config makes the write land early. Elsewhere the measured unsafe cases keep the paste: bash 3.2 runs multi-line lines piecemeal, fish drops an early multi-line write, and any shell loses a line over 1 KB written early.

* test(agent-launch): keep the real-zsh launch-line test out of the Windows lane's gate scan

The Windows lane registration check read `const ZSH_PATH = process.platform === 'win32'` (the head of a multi-line ternary) as a Windows-true flag, so `describe.skipIf(!ZSH_PATH)` looked like a Windows-only suite. The file is POSIX-only; the zsh lookup is now a function.

* refactor(protocol): move the agent.launch capabilities into their own module

Main's protocol-version.ts sits at the 300-line cap, so the prompt-carry capability pushed it over. The four agent.launch capabilities and their doc move to agent-launch-runtime-capability.ts, re-exported by name and spread into RUNTIME_CAPABILITIES at the same position; the advertised lists and every export are unchanged.

* refactor(protocol): import the agent.launch capabilities from their own module

`export *` from protocol-version left the four names undefined under the mobile recording loader, which resolves a relative import through a Proxy with no own keys, so 37 phone recordings lost agent.launchReplay. Importers now name agent-launch-runtime-capability directly; protocol-version only spreads its list.

* refactor(agent-launch): drop the unshipped viewMode field and trusted local caller id

Both were inert in step 1 and existed only for step 2. agent.launch will become a
public plugin API, so every wire field is permanent once shipped; a top-level
viewMode reads as "choose terminal vs chat", which the host decides. Step 2
introduces placement and view intent under a placement object instead.

* fix(agent-launch): read Codex's provisional startup from the rule files' hold anchor

Main (#24375) moved Codex's provisional-header check into codex.json's
provisional_startup hold anchor and deleted codex-terminal-readiness.ts, so the
launch readiness hold now asks showsHoldAnchor, as main's own settled check does.

* fix(agent-launch): hold rule-file name titles to quiet for a launch, and census the zsh fixture

Main (#24375) answers a name-only title from each agent's rule file ahead of the
sustained-title lane, so gemini.json's name_title settled a launch readiness wait
on the shell's auto-title while Gemini was still booting. A launch now asks quiet
of every weak idle verdict, as that lane did.

Main's readiness census requires a recorder for every runtime fixture; the zsh
prompt recording is a non-agent control. Gemini's synthetic baseline is
regenerated for this PR's stated change: a bare gemini title is no longer its
rest mark, so name-only rows settle weak, and a fresh working or blocked status
is no longer overridden.

* fix(agent-launch): paste a launch prompt only when the launched agent is proven in front

A launch pasted its prompt unless a shell was proven in the terminal's
foreground, so any read that could not prove one let the prompt through. After
an agent exited at startup, its shell turned bracketed paste on at the next
prompt, readiness fired on it, and the prompt was typed into the shell:

- macOS: a pane runs its shell under login, so the process-group fence's root
  was never the shell's group and never proved it; the cached foreground name
  could also still name the exited process.
- Windows Git Bash and WSL: the shell-alone-in-its-job check never answers.

Now one fresh read of the terminal's foreground decides: agent, shell or
unknown. Only 'agent' lets a write through (paste, Enter, second Enter, reused
panes too); 'shell' still drops a ready signal. A Windows host never proves the
agent, so there the launch line carries the prompt at any size, as on main.

* test(agent-launch): cover the Windows QA stub, a grok override that exits at once

* fix(agent-launch): keep the local socket alive while a prompted launch waits for its agent

A launch with a prompt now waits up to 60 s for the terminal agent to be
ready before it writes the prompt, and reports not-delivered when the agent
never is. The local runtime socket closes a connection idle for 30 s unless
the request is a long poll, so a launch whose agent exited at startup lost its
reply and the caller saw 'runtime closed the connection' instead of
not-delivered. Classify a prompted agent.launch and agent.launchReplay as a
long poll, as orchestration.workerStart already is for the same wait.

* refactor(agent-launch): narrow the launch params by 'in' instead of a cast

* fix(agent-launch): find a launched agent behind a wrapper that leads its process group

A tcsh or nu launch line runs the agent from /bin/sh '<script>', and a
wrapper script that does not exec its agent does the same: the wrapper leads
the terminal's foreground process group and the agent is a member of it. The
fresh foreground read names the group's leader, sh, so a prompted launch was
refused or pasted late (M4Air tcsh: 2 of 4 not delivered, 2 pasted ~9 s late).

Before that read, take the host's process-group observation as positive proof
when it names the launched agent among the foreground group's members and is
younger than a ready signal's quiet window. It never proves a shell.

* fix(agent-launch): judge the foreground-group proof by when its capture began, not how long ps took

The age the host stamps on a process-group observation runs from the start
of its whole-machine ps, so on a loaded Mac a capture begun after the read
was asked for still read as older than 1 s and the proof was dropped. Count
an observation whose capture began after the read was asked for, less the
window a shared capture is reused across.

* test(agent-launch): keep the crash-guard live test out of the Windows lane's gate scan

The Windows-lane registration scan read the const assigned from a platform
check as a Windows-only gate, though the suite runs everywhere but Windows;
find zsh in a function instead, as the real-zsh typed-line test does.

Under load the fresh foreground scan can fail to answer, which lets the
shell's prompt settle readiness (2 of 4 paired runs). The guard still refuses
that write, so assert the refused write, the property that must always hold.

* perf(agent-launch): read a local pane's foreground from its own terminal, not the whole process table

The foreground read that gates every launch paste ran the daemon's
inspectProcess capture and then a fresh scan, each a whole-machine ps; the
fresh one also waits for any capture already running before it starts its own.
Measured here at load 5: 1.2 s a read (M4Air QA: 3.4-5.0 s, and worker starts
17.6-32 s against main's 9-12 s at load 25-84).

On a local macOS or Linux host, take the pane's root pid from the provider's
session inventory and run one ps limited to that pane's terminal. Its
foreground process group decides: the launched agent or any non-shell member
is the agent (a wrapper that did not exec its agent leads the group), a group
of shells alone is the shell. Same pane, same verdict: 2.7 ms a read. SSH hosts
keep the relay's observation and name.

* test(mobile): re-measure the web app's script sweep after agent.launch's capabilities moved out of protocol-version

The mobile web bundle check failed at 124 assets against a ceiling of 123. Main already sat
exactly on that ceiling: its sweep table read 69 scripts at 16 routes while the tree builds 73,
the whole margin of 4. This branch imports the agent.launch capabilities from their own module,
so protocol-version is no longer pulled into the root layout and four other routes. That moves
which routes share which modules, and the Qoder capability module, imported by protocol-version
and the AI-vault resume path, no longer shares an importer set with anything, so it gets a chunk
of its own: 74 scripts.

The fence says to re-derive the bound rather than raise it, so the sweep is re-measured on this
head (every prefix of the sorted route list). The worst route now adds 10 scripts (session), not
9, which moves the pinned shell crossing from 32 to 30 routes; main re-measured on its own lands
on the same crossing.

* fix(agent-launch): a worker's brief needs its agent found in front, and Grok's start answers on its composer

A paired-server worker start whose agent exited at startup typed its brief into the server's
shell, which ran it: the idle wait can settle on a shell back at its prompt, and the brief was
written with no foreground read. Both worker-start paths now check before each brief write, as a
launch prompt is checked: on a host that can find the agent in front it must be there; on one that
cannot (Windows) a shell proven in front still refuses, and anything else writes as before.

A Grok worker start waited ~10 s more than main: its only rest signal is its bare name, which a
launch holds to quiet output, and Grok animates its logo for ten seconds after its composer glyph.
A worker start for an agent whose rest signal is its bare name and whose composer draws a marker
(Grok, DSH, mimo-code) now also answers on that marker, whichever comes first.
2026-10-05 11:26:52 -07:00

287 lines
13 KiB
JavaScript

import { mkdtemp, rm } from 'node:fs/promises'
import { tmpdir } from 'node:os'
import { join } from 'node:path'
import { fileURLToPath } from 'node:url'
import * as esbuild from 'esbuild'
import { buildMobileWebAppBundle } from './build-mobile-web-app-bundle.mjs'
import { isDirectInvocation } from './script-entry-detection.mjs'
import { assertNoCarriageReturnsInSource } from './mobile-web-source-line-endings.mjs'
import {
MOBILE_WEB_BUNDLE_DIR as defaultBundleDir,
assertMobileWebBundleBuilt
} from './verify-packaged-mobile-web-bundle.cjs'
const projectDir = fileURLToPath(new URL('../..', import.meta.url))
const manifestContract = join(
projectDir,
'src',
'shared',
'mobile-web-bundle',
'manifest-contract.ts'
)
/**
* The document, the route chunks and the images the route tree imports. Derived rather than
* pinned, because a flat number stops agreeing with the chunk ceiling as routes are added: at 128
* and 42 images, 18 routes are already allowed 98 chunks, and 98 + 42 + 1 is 141, so the asset
* count would have failed first and named the count rather than the split that caused it. Written
* as chunks + images + the document, a bundle at the chunk ceiling sits exactly at this one, so
* the chunk ceiling always trips first and the failure says what actually grew.
*/
export function mobileWebAppBundleMaxAssets(routeCount, imageCount) {
return mobileWebAppBundleMaxChunks(routeCount) + imageCount + 1
}
/**
* Phase C byte budget for the app bundle, not the contract ceiling (10 MiB per asset,
* MOBILE_WEB_BUNDLE_MAX_ASSET_BYTES). Deliberately below it so growth trips a build rather than a
* refused asset on a phone. Splitting barely moves it — the same code is emitted in more files —
* so shrinking this still means cutting code.
*
* This head reads 7,686,714 bytes of the 9,437,184 here, 81.5%, leaving 1,750,470. A reading and
* not a pin: nothing asserts it, because the number moves with every build. It is here so the
* generation that spends the rest can see it was already this close.
*/
export const MOBILE_WEB_APP_BUNDLE_MAX_TOTAL_BYTES = 9 * 1024 * 1024
/**
* Scripts emitted at every prefix of the sorted route key list, and the route each prefix added.
*
* A chunk is emitted per distinct set of importers, not per route, so a route's marginal cost is
* what it fails to share rather than what it weighs. Re-measured on this head by building
* `routes.slice(0, n)` for every n, which is what the fence below is derived from rather than
* fitted to. The spread it shows is 1 to 10: `pr` and `web` add one script each, `session` adds ten.
* The root `./_layout.tsx` (the page's web sibling of the native root) sorts first; with it the
* swept tree reads 74 scripts at 16 routes.
*
* This table is the fence's only input, so a route added to the tree stales it and the pins beside
* the fence fail until it is re-measured. That is the point: the bound is re-derived, never bumped.
*/
export const MOBILE_WEB_APP_BUNDLE_SCRIPT_SWEEP = [
['./_layout.tsx', 3],
['./h/[hostId]/[...page].tsx', 7],
['./h/[hostId]/accounts.tsx', 9],
['./h/[hostId]/agent-history/[worktreeId].tsx', 14],
['./h/[hostId]/edit.tsx', 19],
['./h/[hostId]/files/[worktreeId].tsx', 24],
['./h/[hostId]/files/preview/[worktreeId].tsx', 32],
['./h/[hostId]/history/[worktreeId].tsx', 34],
['./h/[hostId]/index.tsx', 39],
['./h/[hostId]/pr/[worktreeId].tsx', 40],
['./h/[hostId]/review/[worktreeId].tsx', 48],
['./h/[hostId]/session/[worktreeId].tsx', 58],
['./h/[hostId]/source-control/[worktreeId].tsx', 63],
['./h/[hostId]/tasks.tsx', 71],
['./h/[hostId]/web.tsx', 72],
['./h/_layout.tsx', 74]
]
const sweptScripts = MOBILE_WEB_APP_BUNDLE_SCRIPT_SWEEP.map(([, scripts]) => scripts)
/** What each route after the first cost, which is the spread the envelope is an upper bound of. */
export const MOBILE_WEB_APP_BUNDLE_ROUTE_SCRIPT_SPREAD = sweptScripts
.slice(1)
.map((scripts, index) => scripts - sweptScripts[index])
/**
* How far above the measurement the envelope sits, and the only slack a refactor gets.
*
* Measured, not chosen, and per head rather than cumulative: at 8, 10, 12 and 14 routes the head
* that wrote the old `4r + 16` read 32, 43, 61 and 69, the head that first swept these prefixes
* read 34, 44, 57 and 65, and this one reads 33, 43, 56 and 64. So one head has moved the count by
* as much as four at a fixed route count with no route added (61 to 57), and the step that dropped
* the page's second Zod moved it by one everywhere. Four is that worst step, which is what a shared
* importer set moving between heads costs. Summing the steps instead would grow this number every
* head and loosen the fence for free. A refactor inside four keeps building; anything past it
* re-measures the sweep.
*/
export const MOBILE_WEB_APP_BUNDLE_SCRIPT_MARGIN = 4
/**
* How many scripts the page may be cut into, for a given number of routes.
*
* Anchored on the sweep above rather than fitted to the route count, because the count is not a
* function of the route count alone: `4r + 16` was a guess at break-even and its slack ran from 17
* at one route to 4 at thirteen, so it was a near-miss exactly where the tree actually sits. This
* is the measurement plus the margin at the swept tree, growing by the worst route the sweep saw
* for every route past it — a route cannot breach it without costing more than any route measured.
*
* Flat below the swept length, which is the whole sweep's upper bound too: the count rises with the
* prefix, so one number above its top row is above every row. The fence is only ever asked about
* the real tree, and routes are only ever added.
*
* Two-sided in the test beside it. An envelope more than the margin above the build is a fence
* nobody re-derived, and it fails there rather than surviving as headroom for a bump.
*
* The route count stays the only term. A deferred engine belongs inside one artifact and costs one
* script: C7.10 item B first reached mermaid with `import('mermaid')`, which emitted 103 more
* because mermaid lazily imports each of its own diagram types, and a second term admitting those
* would have raised this fence far enough to admit any split at all. The build test's control is
* what holds that line.
*
* This is the ceiling that catches a split running away; MOBILE_WEB_APP_BUNDLE_MAX_ENTRY_BYTES
* below is the one that catches it collapsing, and it is the real budget of the two.
*/
export function mobileWebAppBundleMaxChunks(routeCount) {
const beyondTheSweep = Math.max(0, routeCount - MOBILE_WEB_APP_BUNDLE_SCRIPT_SWEEP.length)
return (
sweptScripts.at(-1) +
MOBILE_WEB_APP_BUNDLE_SCRIPT_MARGIN +
Math.max(...MOBILE_WEB_APP_BUNDLE_ROUTE_SCRIPT_SPREAD) * beyondTheSweep
)
}
/**
* What the browser must parse before the first route can paint: the entry plus every chunk it
* reaches by static import. This is the budget splitting exists to hold — it was 8.16 MB as one
* chunk and measures 1,244,312 bytes split on this head, 1.19 of the 3 MiB — so a route
* re-imported statically, or `splitting` dropped, fails the build here instead of arriving as a
* slow first open on a phone.
*
* It is not a per-route escape hatch. Re-measured here by making one route's manifest entry a
* static import and reading this same closure back: session alone breaks the bound at 3.32 MiB,
* and tasks at 2.17, source-control 2.04, review 2.03, index 1.89 and files/preview 1.85 each
* spend most of a budget that has to cover the entry as well. What keeps the hatch usable at all
* is that expo-router reads `unstable_settings` off layout nodes only, and the subtree's one
* layout, `h/_layout.tsx`, measures 1.89 MiB static. Any other route needing a synchronous export
* needs this number re-measured, not a static import.
*/
export const MOBILE_WEB_APP_BUNDLE_MAX_ENTRY_BYTES = 3 * 1024 * 1024
/** Every tree whose bytes reach the buildId, so a CRLF checkout cannot fork it. */
export const MOBILE_WEB_APP_SOURCE_DIRS = [
join(projectDir, 'mobile', 'web-entry'),
join(projectDir, 'mobile', 'app'),
join(projectDir, 'mobile', 'src')
]
class VerificationError extends Error {}
function fail(message) {
throw new VerificationError(message)
}
/**
* How many assets the phone will accept, read from the contract rather than copied: the native
* shells hold their own 256 and refuse a larger manifest outright. Bundled through esbuild
* because node cannot resolve that module's extensionless TypeScript imports, so the number is
* evaluated from the contract and not parsed out of it.
*/
export async function readMobileWebBundleMaxAssets() {
const { outputFiles } = await esbuild.build({
entryPoints: [manifestContract],
bundle: true,
write: false,
format: 'esm',
platform: 'node',
logLevel: 'silent'
})
const source = Buffer.from(outputFiles[0].contents).toString('base64')
const { MOBILE_WEB_BUNDLE_MAX_ASSETS: ceiling } = await import(
`data:text/javascript;base64,${source}`
)
if (typeof ceiling !== 'number') {
fail(`${manifestContract} exports no MOBILE_WEB_BUNDLE_MAX_ASSETS to bound the build with`)
}
return ceiling
}
/**
* The derived ceiling is only a budget while it stays inside the map the phone can hold: the
* shells return null for a manifest over MOBILE_WEB_BUNDLE_MAX_ASSETS rather than dropping the
* extra assets, so a route count that pushes the chunk envelope plus images plus the document past
* it would pass this build and fail on the device with nothing to read. At today's 42 images that
* is 29 routes, inside what Phase C adds, which is why this is a build failure and not a comment.
* The envelope grants the worst swept route to each one past the sweep, so re-measuring a tree
* whose routes share more moves that crossing out again.
*/
export function assertAssetCeilingFitsShell(routeCount, imageCount, shellMaxAssets) {
const ceiling = mobileWebAppBundleMaxAssets(routeCount, imageCount)
if (ceiling > shellMaxAssets) {
fail(
`the ceiling derived for ${String(routeCount)} route(s) and ${String(imageCount)} image(s) ` +
`is ${String(ceiling)} assets, over the ${String(shellMaxAssets)} the shell will load`
)
}
return ceiling
}
async function buildIntoScratch() {
const scratch = await mkdtemp(join(tmpdir(), 'orca-mobile-web-app-verify-'))
try {
return await buildMobileWebAppBundle({ outDir: join(scratch, 'mobile-web') })
} finally {
await rm(scratch, { recursive: true, force: true })
}
}
// bundleDir is a seam for the tests, which verify a scratch build; the script always verifies out/.
export async function verifyMobileWebAppBundle({ bundleDir = defaultBundleDir } = {}) {
for (const directory of MOBILE_WEB_APP_SOURCE_DIRS) {
await assertNoCarriageReturnsInSource(directory)
}
const manifest = assertMobileWebBundleBuilt(bundleDir)
if (manifest.totalBytes > MOBILE_WEB_APP_BUNDLE_MAX_TOTAL_BYTES) {
fail(
`bundle is ${String(manifest.totalBytes)} bytes, over the Phase C budget of ` +
`${String(MOBILE_WEB_APP_BUNDLE_MAX_TOTAL_BYTES)}`
)
}
const first = await buildIntoScratch()
const second = await buildIntoScratch()
if (first.manifest.buildId !== second.manifest.buildId) {
fail(`buildId is not reproducible: ${first.manifest.buildId} then ${second.manifest.buildId}`)
}
if (first.manifest.buildId !== manifest.buildId) {
fail(
`${bundleDir} is stale: it carries buildId ${manifest.buildId}, a fresh build produces ${first.manifest.buildId}`
)
}
// Read off the fresh build rather than the manifest: neither bound is a manifest field, and the
// buildId just proved this build is the one on disk.
// After the fresh build, which is what knows how many of the assets are images.
const maxAssets = assertAssetCeilingFitsShell(
first.routeKeys.length,
first.imageCount,
await readMobileWebBundleMaxAssets()
)
if (manifest.assets.length > maxAssets) {
fail(
`bundle has ${String(manifest.assets.length)} assets, over the Phase C budget of ` +
`${String(maxAssets)} for ${String(first.routeKeys.length)} route(s) and ` +
`${String(first.imageCount)} image(s)`
)
}
const maxChunks = mobileWebAppBundleMaxChunks(first.routeKeys.length)
if (first.chunkCount > maxChunks) {
fail(
`bundle is cut into ${String(first.chunkCount)} chunks, over the Phase C budget of ` +
`${String(maxChunks)} for ${String(first.routeKeys.length)} route(s)`
)
}
if (first.entryStaticBytes > MOBILE_WEB_APP_BUNDLE_MAX_ENTRY_BYTES) {
fail(
`${String(first.entryStaticBytes)} bytes load before the first route, over the Phase C ` +
`budget of ${String(MOBILE_WEB_APP_BUNDLE_MAX_ENTRY_BYTES)}`
)
}
return manifest
}
if (isDirectInvocation(import.meta.url, process.argv[1])) {
try {
const manifest = await verifyMobileWebAppBundle()
console.log(
`[verify-mobile-web-app-bundle] OK — ${String(manifest.assets.length)} asset(s), ` +
`${String(manifest.totalBytes)}/${String(MOBILE_WEB_APP_BUNDLE_MAX_TOTAL_BYTES)} bytes, ` +
`reproducible buildId ${manifest.buildId}`
)
} catch (error) {
console.error(`[verify-mobile-web-app-bundle] ${error.message}`)
process.exit(1)
}
}