mirror of
https://github.com/stablyai/orca.git
synced 2026-09-26 08:02:38 +00:00
* feat(codex): backfill managed-home sessions into the real Codex home once per host Orca-launched Codex sessions currently land only in the Orca-managed runtime home, so the user's own `codex resume` picker and app history never see them (#4444, #8612). Backfill the managed sessions tree into the real ~/.codex/sessions/YYYY/MM/DD layout once per host: - hardlink first (one physical rollout log), copy as the cross-volume fallback; existing target files are always skipped, nothing in either home is deleted or moved - idempotent; per-file failures leave the completion marker unset so the next startup retries cheaply - JSONL audit log of every link/copy/failure under <userData>/codex-session-backfill/ - honors the custom Codex session source home override, mirroring the existing system->managed bridge WSL managed homes are distro-local and need an in-distro variant; that is a follow-up. * feat(codex): flag-gated system-default real-home routing scaffolding Staged internal flag (default OFF, no settings UI): route the SYSTEM-DEFAULT Codex account at the user's real ~/.codex instead of Orca's managed runtime home. Flag OFF is byte-identical to today; managed (multi-account) selections are unchanged in either state. Routing (flag ON + host system default = no managed account): - CodexRuntimeHomeService.prepareForCodexLaunch / prepareForRateLimitFetch return null so the PTY/env layer injects no managed CODEX_HOME and the rate-limit fetcher + auth-presence gate fall back to ~/.codex (the background poller stops spawning Codex against the managed home — the #5370 auth war). - buildPtyHostEnv strips only a nested-Orca-inherited Orca-owned override (CODEX_HOME matching the private ORCA_CODEX_HOME marker), preserving a user-set CODEX_HOME. Shell-ready re-exports already no-op without the marker. - The headless commit-message Codex path strips the same inherited override. Hook install for the real-home lane (append-last into ~/.codex/hooks.json, trust via the app-server client) lands with the trust plumbing; the managed hook install is skipped for this lane meanwhile. Credit @jellychoco (#8606) for the native-home routing direction. Depends on the codex trust-rpc-grant plumbing for the real-home hook installer. * fix(codex): strip the daemon-inherited Orca CODEX_HOME override for real-home routing The daemon spawns PTYs from its own inherited environment and honors only spawnOptions.envToDelete, so mutating the sparse env object was not enough to strip an Orca-owned CODEX_HOME the daemon already carries. Add the strip to envToDelete for both daemon host-spawn paths, preserving a user-set CODEX_HOME. Verified live via CDP against a sandboxed dev instance (flag ON): an Orca-spawned pane reports empty CODEX_HOME/ORCA_CODEX_HOME, so Codex resolves its own ~/.codex. Adds daemon-path unit coverage (strip Orca-owned, preserve user-owned, no-op when flag OFF). * fix(codex): harden one-time session backfill * test(codex): cover staged cross-volume install * feat(codex): app-server trust-grant client, capability cache, and grant ledger Short-lived codex app-server JSON-RPC client (hooks/list + config/batchWrite, the same pair the Codex TUI 'Trust all' flow calls), run in a bundled ELECTRON_RUN_AS_NODE entry so synchronous launch prep can block on it with a hard deadline and guaranteed child reap. Capability cache modeled on GitCapabilityCache, scoped per execution host (native vs each WSL distro), with a narrow unknown-method/missing-subcommand unsupported predicate. The grant ledger records verified grants so steady-state launches skip the RPC. * fix(codex): grant managed hook trust via codex app-server RPCs in install/refresh Host and WSL installs now grant trust for Orca's managed status hooks through codex's own hooks/list -> config/batchWrite -> re-list verify, scoped to exactly the managed entries; the previous computeTrustedHash lane is the unchanged fallback for incapable/erroring CLIs. getStatus and the removal paths recognize ledger-recorded codex hashes so drift between codex's real algorithm and the replica no longer misreports or strands trust. SSH remote install is untouched by design. * test(codex): cover app-server trust grant client, cache, ledger, and lanes * test(codex): cover commit-message real-home override strip/preserve Adds the two cases for the headless commit-message Codex env under real-home routing: a nested-Orca-inherited Orca-owned CODEX_HOME is stripped, and a user-owned CODEX_HOME is preserved. * test(codex): WSL grant-lane coverage — in-distro invocation and fallback parity * feat(codex): real-home hook installer trusted via the codex app-server grant client With the real-home flag ON and the system-default selection, install Orca's status hook into the user's real ~/.codex before any pane spawns: - entry APPENDED LAST per managed event: codex hook trust keys are positional (source:event:group:handler), so appending keeps every user entry's position and trust record intact; user entries and unknown top-level hooks.json fields are preserved verbatim - trust is granted exclusively through the codex app-server client (hooks/list + config/batchWrite, verified by re-list); Orca never writes [hooks.state] into the user's real config.toml itself - if the grant lane is unavailable (old binary, unsupported RPC, verify failure), the appended entry is rolled back byte-exactly and the host keeps the managed-home lane end to end (PTY env, rate limits, commit messages) via a lane gate on the runtime-home service - one-time pristine backup of the user's hooks.json under Orca's userData; a rolling .bak sits next to the file (existing atomic writer) - hook opt-out sweeps Orca entries from the real home and drops Orca-owned trust records; flag-off downgrade re-arms the existing legacy system-home sweep, which removes the entry and its trust keys cleanly - the legacy system-home sweep is suppressed only while the real-home lane owns ~/.codex/hooks.json, so managed installs cannot delete the entry * fix(codex): resolve the trust-grant entry without requiring electron The grant bridge is reachable from plain-Node CLI entries, where the plain-node entry guard rejects any chunk containing require("electron"). Resolve the bundled session entry from __dirname (root chunk and chunks/ layouts) with an app.asar -> app.asar.unpacked rewrite for packaged runs, instead of electron's app path APIs. * fix(codex): keep session backfill off main thread Use asynchronous, sequential filesystem operations for the one-time rollout backfill, and avoid repeated target-directory probes. Treat inaccessible managed session roots as retryable failures instead of writing a false completion marker. * fix(codex): harden app-server trust grant fallback * fix(codex): install cross-volume session backfill copies atomically On a real Codex home whose filesystem supports no hardlinks (exFAT/FAT, some network mounts), the staged cross-volume copy was installed with a non-atomic copyFile(..., COPYFILE_EXCL) straight into the final rollout-*.jsonl name. An install interrupted mid-copy (app quit, crash, ENOSPC during the deferred run) could strand a truncated rollout that the next run then skips as already-present, defeating the staging design's own guarantee that a failed copy never leaves a partial session behind. Install the fully-staged copy with an atomic rename instead, guarded by an existence re-check so it keeps the never-overwrite contract (and the rename source is the same immutable managed rollout, so any clobber would be byte-identical). Cover the no-hardlink-support target and an interrupted install that must leave no partial in the user's sessions tree. * fix(codex): resolve grant entry from __dirname so plain-node CLI entries stay electron-free The build guard rejects any electron require reachable from plain-node entries; the bridge now maps app.asar to app.asar.unpacked by string replacement instead of consulting electron app paths. CLI typecheck project lists the new trust-grant module graph. * fix(codex): harden trust grant reconciliation * fix(codex): restore trust config permissions on rollback * fix(codex): harden real-home routing cleanup and retries * fix(codex): preserve unicode trust RPC responses * fix(codex): preserve remote env and complete real-home cleanup * fix(codex): preserve real-home lane invariants * test(terminal): isolate replacement idle reset assertion * fix(codex): preserve real-home dotfile links * fix(codex): preserve verified trust grants across launch prep * fix(codex): preserve dangling config symlinks on rollback * fix(codex): don't revoke a just-granted WSL home on a false 'missing' probe The async wsl.exe canonical-path settlement could report the runtime home 'missing' immediately after a verified RPC grant (a false negative — codex had just written and re-listed trust there), which drove the reconciliation 'remove' branch to delete all six granted [hooks.state] tables, leaving a bare [hooks.state] the launching pane read as 'hooks need review'. A 'missing' settlement now revokes only when no successful install ran this generation; a genuinely moved home still resolves to a different path and reinstalls. * test(codex): model codex config/batchWrite faithfully on Windows The grant-lane stub simulated codex by calling Orca's upsertHookTrustEntries, which writes both separator variants for a Windows key (a fallback-lane compat shim real codex never does) — fabricating duplicate tables and whitespace the RPC path never produces, so the byte-stable and no-duplicate assertions failed on win32. Replace it with a single-variant, blank-line-separated writer that matches the real 0.144.x binary's output. * feat(codex): collapse duplicate session listings across Codex roots Backfilled/bridged rollouts are hardlinked into both the real ~/.codex and Orca's managed runtime home, so AI Vault listed each session once per root (#7521). Dedup candidates by rollout file name pre-parse and parsed sessions by session id post-parse, keeping the canonical root: host real home first (unprefixed resume), then the managed runtime home, then other homes. Applies to local, WSL, and SSH-remote scans. * feat(codex): background sqlite index heal for backfilled sessions Codex's own state-DB metadata backfill is one-shot, so rollouts hardlinked in by Orca's session backfill never become visible to Codex's DB-driven surfaces. Extract the app-server stdio JSONL transport into codex-app-server-session (shared with the trust-grant client) and add a bounded, resumable background pass that drives Codex's lazy indexing via thread/read per backfilled session: recent-first, batched onto one short-lived server per batch with small concurrency, ledger + marker so steady-state startups are a no-op, stop-aware on quit, and capability-aware on CLIs without the app-server surface. * fix(codex): preserve session identity during dedup heal * fix(codex): preserve user trust during real-home cleanup * fix(codex): harden real-home heal boundaries * fix(codex): fail closed on unsafe backfill install * fix: harden real-home hook cleanup * fix(ai-vault): preserve execution boundaries and reap children * fix(codex): narrow app-server unsupported detection * fix(codex): bound user hook trust rebase retries per host The rebase lane ran a codex app-server session on every launch prep while a host was stuck (CLI without app-server support, or keys hooks/list cannot match). Gate the transaction on the shared capability cache and add the same 5-minute transient cooldown the grant lane uses, so sweep and legacy-cleanup retries cost plain fs reads instead of a codex session per pane spawn. * fix(codex): enforce real-home resume and heal boundaries * fix(codex): establish real-home lane before cleanup * fix(codex): stop index heal before delayed spawn * fix(codex): protect symlinked rolling backups * fix(ai-vault): preserve resume env deletion through drag * fix(codex): strip inherited Codex homes on mobile real-home resume The mobile resume surface types a bare real-home codex resume into a freshly created pane, but never asked for CODEX_HOME/ORCA_CODEX_HOME deletion at pane spawn, so an agentDefaultEnv-pinned or daemon-inherited Codex home rerouted the resume away from the user's real ~/.codex while the same session resumed correctly on desktop. Share the deletion helper from the AI Vault resume builders and forward it through the mobile launch and session.tabs.createTerminal call. * fix(codex): gate session migration on real-home lane * fix(codex): stop session backfill after opt-out * fix(codex): keep session heal failures retryable * fix(codex): keep session migration state recoverable * fix(codex): retry republished missing session heals * fix(codex): preserve hook symlink trust path * fix(codex): disambiguate POSIX trust paths * fix(codex): align hook trust source paths * fix(codex): harden trust grant lifecycle * fix(codex): restore envToDelete on client invocation type after base reconcile * test(codex): type child.stdout as PassThrough for oversized-output write * Assemble RC: reconcile app-server transport API across PRs Unify on the object RPC surface from the index-heal transport (#8921) while preserving the default-home env strip (#8828) and the narrowed missing-app-server capability signal (#8847): adapt the user-hook-trust-rebase consumer + tests, port envToDelete stripping into the shared session, and route stderr classification through the canonical capability-signal module. * RC: enable system-default real-home routing by default (flag ON) Flip codexSystemDefaultRealHomeEnabled to default ON for this RC's staged rollout (a user can still opt out by setting it false, which stays byte-identical to managed-home behavior). This is the only intended behavior difference between the RC branch and the individual PRs. Updates the two tests that assumed the prior OFF default. * fix(codex): snapshot hooks.json bytes+parse in one read to close real-home clobber race The install/sweep/legacy-cleanup paths parsed hooks.json, then did a separate later read to capture the previous bytes for the pre-write generation guard. A concurrent save (second Orca instance or the user editing the file) could land between the parse and that second read and be silently overwritten. readHooksJsonWithRaw returns the raw bytes and parse from a single read so the guard compares against exactly what it parsed. Adds a regression test that mutates hooks.json mid-RPC and asserts the sweep aborts without clobbering. * fix(codex): sanitize managed account config trust * fix(codex): guard OAuth add for custom providers * fix(codex): persist outgoing managed tokens before real-home lane takeover (PR-C) prepareForCodexLaunch returns null early for the real-home / system-default lane before syncForCurrentSelection runs. If a managed account is still recorded as synced when the selection has dropped to the system default (nulled without a sync pass, or auto-deselect on missing managed auth), a Codex-refreshed token stranded in the shared runtime home is never persisted to its canonical per-account home -> token loss. Read the outgoing managed account's refreshed token back before the real home takes over. The real-home lane implies host === null, so running the managed->system-default transition restores only Orca's runtime mirror from ~/.codex and never writes the real ~/.codex. It is a no-op once the selection has already been reconciled, so the normal select path does not double-write. * fix(codex): preserve refreshes across all default transitions * feat(codex): show system-default/real-home account identity in switcher (PR-B) The account switcher modeled the system-default Codex account as activeAccountId:null with no identity fields, so the null row rendered blank ("System default" / generic subtitle) even though its effective login is whatever ~/.codex/auth.json currently is. Add a CodexSystemDefaultIdentity descriptor {hasAuth, authKind, email, providerAccountId, workspaceLabel} to CodexRateLimitAccountsState, resolved live and READ-ONLY from ~/.codex by the accounts service and returned from listAccounts()/getSnapshot(). The settings switcher now renders the null (system-default) row as that real identity: the OAuth email when signed in, "Custom provider — no usage tracked." for env-key/custom-provider logins (auth.json with OPENAI_API_KEY, or an OPENAI_API_KEY env with no auth.json), and the generic fallback when signed out. Identity is host-scoped (per-distro WSL keeps the generic label). Orca never writes ~/.codex; managed-account switches only touch Orca-owned homes, so the system-default identity stays a stable, displayed source of truth. Usage already routes to the real home via getSystemCodexHomePath, so the switcher now attributes it to a real face. Tests (sandboxed temp homes only): OAuth email/provider resolution, api-key auth.json and env-key (no auth.json) as custom-provider, signed-out, and select/deselect of a managed account never mutating ~/.codex/auth.json. * fix(codex): parse multiline provider pins in OAuth guard * fix(codex): harden managed trust sanitization * fix(codex): harden system-default identity rendering * feat(codex): give each managed account a self-contained CODEX_HOME; retire shared mirror (PR-E) With the real-home flag ON, a host managed account now launches directly against its own codex-accounts/<id>/home instead of the shared runtime mirror + auth.json hot-swap: - codex-home-paths: syncSystemCodexResourcesIntoManagedHome links system resources into any managed home (ownership-marker discipline; never symlinks into / mutates ~/.codex). - runtime-home-service: prepareForCodexLaunch / prepareForRateLimitFetch / syncForCurrentSelection route the per-account home directly and skip the shared-home hot-swap + token read-back; each home keeps its own auth in place (fixes GAP-5 concurrent auth race). Session discovery scans every per-account home. - hook-service / hook-trust-promotion: install/getStatus/refresh accept a runtimeHomePath so hooks + RPC-granted trust land in the per-account home. - service: config mirror into a self-contained home uses the trust- preserving merge so granted hook/project trust survives account switches. - codex-session-root-dedup: rank codex-accounts/<id>/home as canonical managed alongside the shared runtime home. Flag-OFF and the system-default real-home (null) lane are unchanged; the nested-Orca CODEX_HOME===ORCA_CODEX_HOME daemon strip (#5370) is preserved. Sandboxed tests only; ~/.codex is never mutated. * fix(codex): validate per-account home ownership * fix(codex): keep managed rollouts discoverable across real-home opt-out WI-4 lossless migration/rollback validation for pre-E shared-mirror managed accounts. Session discovery gated the per-account home scan on the real-home flag, so opting back out (flag OFF) hid every rollout an account accumulated while the flag was ON — the data stayed on disk but vanished from the AI Vault until the flag flipped back on. Scan a managed host home whenever it holds a sessions/ tree, independent of the flag; a never-enabled install keeps its homes credential-only so opt-out stays byte-identical to today. Forward migration was already lossless (the shared mirror is always scanned) and the opt-out credential read-back already refuses to overwrite a fresher per-account token; add tests locking all three invariants. Sandboxed tests only; ~/.codex is never touched. * fix(codex): migrate stranded shared auth on E takeover * test(e2e): isolate Electron from developer Codex home * test(codex): add real-account validation harness * fix(codex): finish C and E matcher composition * fix(codex): bound validation harness shutdown * test(codex): isolate hook lifecycle user data * test(codex): cover realistic account-home migration * fix(codex): keep standalone home tripwire active * test(codex): fingerprint system auth in validation reports * fix(codex): bind managed homes to account ownership * fix(codex): normalize Windows trust source identity * fix(codex): make Windows trust upgrade transactional * test(codex): use TypeScript pipeline for validation scripts * test(codex): run validation modules through native node * test(codex): allow slow Windows tripwire startup * fix(codex): survive lingering Windows codex login processes in add-account On Windows, codex login can keep running (with descendants) after it has written auth.json, holding OS handles on the per-account managed home (log/codex-login.log). That made doAddAccount's post-login cleanup fail with ENOTEMPTY (rmSync) and left an orphaned codex-accounts/<id>/home. - runCodexLogin now watches for auth.json on Windows and force-kills the login process tree (taskkill /t) if it lingers past a short grace period; the forced exit is treated as a successful login. The 120s timeout path also kills the whole tree instead of only the direct child. macOS/Linux behavior is unchanged. - safeRemoveManagedHome now removes homes with rmSync maxRetries / retryDelay (mirroring the local-worktree-filesystem Windows policy) and no longer lets a cleanup failure mask the original add error. - run-codex-real-account-validation.mjs accepts --temp-parent / ORCA_CODEX_VALIDATION_TEMP_PARENT so the disposable root can live outside %USERPROFILE% on Windows, and fails with an actionable message before creating anything when the temp parent is inside the primary home. The real-home guard is unchanged. * fix(codex): preserve managed-account MCP .credentials.json on per-account-home migration (#8440) Codex file-mode MCP OAuth tokens live in $CODEX_HOME/.credentials.json, keyed by MCP server URL with no account identity of their own. The legacy shared-mirror -> per-account-home migration only carried auth.json, so an existing managed account with authed MCP servers had its tokens stranded on upgrade and silently needed re-auth. Carry the shared mirror's .credentials.json into the same identity-proven per-account home alongside auth.json: only into the single uniquely-matched active account (no cross-account leak), only when the destination has none yet (never clobber a newer file the account authed in its own home), atomic 0600, absent-source no-op. New MCP auth already lands in the per-account home since that home is CODEX_HOME. * fix(codex): preserve Windows reauthentication login flow * test(codex): build real-account validation harness cross-platform on Windows The harness built its app with execFileSync('npx', ['electron-vite', ...]), but npx resolves to a .cmd shim on Windows that execFileSync cannot launch (ENOENT), so the harness could not build its own app there and required --skip-build with a prebuilt out/main/index.js. Extract resolveElectronViteBuildCommand(repoRoot): it runs the repository-local electron-vite JS entry (node_modules/electron-vite/bin/electron-vite.js) with the current Node binary (process.execPath), which resolves identically on macOS, Linux, and Windows with no shell. It throws a clear error if the local entry is missing (install deps or pass --skip-build). --skip-build behavior is unchanged. Add regression coverage asserting the build command uses process.execPath and the repo-local JS entry (not npx), and that a missing entry fails clearly. * fix(codex): version the MCP creds migration independently of the auth marker The auth carry and the MCP .credentials.json carry (#8440) shared one existence-only v1 marker, so any build that stamped the auth-only marker first would strand the MCP store forever. The MCP carry now concludes via its own per-account-mcp-creds-migration-v1.json marker and runs even when the auth marker is already present; ordering is code-enforced instead of landing-discipline-enforced. Also isolate per-account read failures: one stale or deleted account home no longer aborts the whole migration. The broken account stays in the unique-identity ambiguity gate via its stored fields but is never read or written, so the active account still migrates. * fix(codex): fail corrupt managed auth.json without echoing credential bytes A raw JSON.parse SyntaxError from loadOAuthCredentials could carry auth file fragments into logs and the add/reauth error surface. Throw a sanitized error instead; filesystem errors still propagate unchanged. * fix(mobile): give the pairing runtime a disposable home for the E2E boot guard The main-process guard now refuses to start with ORCA_E2E_USER_DATA_DIR set but the real user home, and this was the one caller not updated — the temporary pairing runtime crashed before emitting its pairing URL. * test(codex): canonicalize harness containment guards and retry cleanup Resolve symlinks before the disposable-root containment checks so a symlinked temp parent cannot smuggle the throwaway home inside the primary home, and give the final cleanup rm Windows retry/force so a briefly lingering codex handle cannot strand the credential-bearing root. * test(codex): add lane-aware containment mode to the real-account harness The Windows gate-D run proved strict zero-event whole-profile containment is structurally unreachable with the real-home flag ON: system-default spawn sites deliberately delete CODEX_HOME so native codex resolves the real ~/.codex, and on Windows the binary ignores the USERPROFILE sandbox. Its own volatile runtime churn (root sqlite/WAL/SHM, tmp/, log/) is the shipped Phase-1 design, not a candidate defect. --lane-aware-containment records those designed events without aborting while every other real-home write — auth.json, config.toml, .credentials.json, hooks.json, sessions/, anything unknown — remains a hard violation and still aborts the run. Default behavior is unchanged (strict); the absolute zero-event claim stays carried by macOS runs, where HOME does sandbox native codex. * test(codex): allow the real-account harness to pin the real-home flag off --system-default-real-home off seeds and env-pins the flag OFF so every codex spawn gets an explicit managed CODEX_HOME and native codex never resolves the OS profile. This is the only Windows configuration where the strict zero-event whole-profile tripwire is reachable, and it matches the stable-rollout default; flag-ON runs keep lane-aware classification. * test(codex): correct the flag-off harness comment to kill-switch rationale The rollout ships all codex-home changes at once (no phased rollout), so flag OFF is the emergency kill-switch lane, not the stable default. * test(e2e): canonicalize the isolated E2E home path The disposable HOME lives under os.tmpdir(), whose spelling is an alias on CI (macOS /var symlink, Windows 8.3 RUNNER~1). Git canonicalizes worktree paths, so worktrees created under the aliased home never matched the app's listing — golden core flows and the packaged crash-survival harness failed with 'worktree created but not found in listing'. Resolve the home to its canonical spelling at creation in both the e2e helper and the packaged-app driver. * fix(codex): address CodeRabbit review on the landing PR - carry envToDelete through the mobile agent-resume startup plan so a real-home Codex resume cannot inherit an ambient CODEX_HOME - strip Orca-owned Codex overrides in the commit-message WSL fallback, matching the host fallback - strip ELECTRON_RUN_AS_NODE in the computer-e2e driver like every other home-isolation caller - drop the unused hooksEnabled parameter from isRealHomeCodexHookLaneUsable * feat(codex): ship real-home routing unconditionally, remove the rollout flag The codexSystemDefaultRealHomeEnabled setting is gone from types and constants and the helper no longer consults settings — the system-default real-home lane and per-account homes ship for everyone in one release. This also un-strands profiles that rc-era builds stamped with false (the setting had no UI, so every stored false was a seeded artifact that would have silently kept those users on the legacy mirror forever). The ORCA_CODEX_SYSTEM_DEFAULT_REAL_HOME env override survives strictly as a test-rig control: the containment harness pins the legacy lane for strict zero-event Windows runs, e2e home isolation pins lanes inside disposable homes, and the legacy-lane test suites now route their per-test lane selection through it. --------- Co-authored-by: OrcaWin <alpha-eng@stably.ai>
480 lines
18 KiB
JavaScript
480 lines
18 KiB
JavaScript
// Drive the installed, packaged Orca app with Playwright's Electron driver.
|
|
//
|
|
// This targets a PRODUCTION build, so it must NOT depend on the e2e-only store
|
|
// exposure (window.__store / window.__paneManagers) — those exist only under a
|
|
// `--mode e2e` / VITE_EXPOSE_STORE build. Everything here uses ARIA/DOM
|
|
// selectors that ship in production (matching tests/e2e/helpers/terminal.ts and
|
|
// terminal-attention.spec.ts) and proves interactivity through filesystem
|
|
// sentinels rather than by reading the WebGL-rendered xterm buffer:
|
|
// - typed commands write a marker FILE; the harness checks the file. This
|
|
// proves keystrokes reached the shell AND executed — stronger, and robust,
|
|
// than scraping canvas-rendered terminal text.
|
|
//
|
|
// The long-running marker also sets a unique window-title canary and writes a
|
|
// heartbeat file every ~500ms: the canary lets the window watch attribute any
|
|
// real console flash to our child, and the heartbeat proves the session is
|
|
// live and streaming.
|
|
|
|
import { _electron as electron } from '@stablyai/playwright-test'
|
|
import { execFileSync } from 'node:child_process'
|
|
import { mkdirSync, realpathSync, writeFileSync } from 'node:fs'
|
|
import path from 'node:path'
|
|
import { seedFreshProfile } from './onboarding-profile.mjs'
|
|
|
|
const NEW_TAB_BUTTON = { role: 'button', name: 'New tab' }
|
|
const NEW_TERMINAL_ITEM = /New Terminal/i
|
|
const NEW_WORKSPACE_BUTTON = { role: 'button', name: 'New workspace' }
|
|
const SORTABLE_TAB = '[data-testid="sortable-tab"]'
|
|
// Why: the layout mounts hidden duplicate panes; only the visible one is the
|
|
// live terminal, so target `:visible` to avoid focusing/measuring a hidden copy.
|
|
const TERMINAL_SURFACE_VISIBLE = '[data-terminal-tab-id]:visible'
|
|
const XTERM_CONTAINER_VISIBLE = '.xterm:visible'
|
|
const XTERM_INPUT = '.xterm-helper-textarea'
|
|
const RESTRICTED_E2E_ENV_KEYS = new Set([
|
|
'HOME',
|
|
'USERPROFILE',
|
|
'CODEX_HOME',
|
|
'ORCA_CODEX_HOME',
|
|
'ORCA_CODEX_SYSTEM_DEFAULT_REAL_HOME',
|
|
'ORCA_E2E_HOME_DIR',
|
|
'ORCA_E2E_USER_DATA_DIR'
|
|
])
|
|
|
|
/**
|
|
* Launch the installed Orca.exe. Pointing userDataDir at a harness-owned temp
|
|
* dir isolates this run's daemon (its socket/token path becomes unique), so
|
|
* daemon lookups never collide with other Orca installs/daemons on the box.
|
|
* Pass `seedProfile` (a buildFreshProfile object) to write orca-data.json
|
|
* BEFORE this launch — do so only on the FIRST launch, never before the
|
|
* post-update relaunch, or the persisted session under test is destroyed.
|
|
*/
|
|
export async function launchInstalledApp({
|
|
exePath,
|
|
userDataDir,
|
|
seedProfile = null,
|
|
extraEnv = {}
|
|
}) {
|
|
const {
|
|
ELECTRON_RUN_AS_NODE: _drop,
|
|
CODEX_HOME: _codexHome,
|
|
ORCA_CODEX_HOME: _orcaCodexHome,
|
|
...cleanEnv
|
|
} = process.env
|
|
void _drop
|
|
void _codexHome
|
|
void _orcaCodexHome
|
|
const restrictedExtraEnvKey = Object.keys(extraEnv).find((key) =>
|
|
RESTRICTED_E2E_ENV_KEYS.has(key.toUpperCase())
|
|
)
|
|
if (restrictedExtraEnvKey) {
|
|
throw new Error(`extraEnv.${restrictedExtraEnvKey} cannot override E2E home isolation`)
|
|
}
|
|
mkdirSync(userDataDir, { recursive: true })
|
|
if (seedProfile) {
|
|
seedFreshProfile(userDataDir, seedProfile)
|
|
}
|
|
// Why: userData relocation does not change Node's home; the packaged E2E
|
|
// must not resolve the default Codex account against the runner's profile.
|
|
const requestedIsolatedHome = path.join(userDataDir, 'home')
|
|
mkdirSync(requestedIsolatedHome, { recursive: true })
|
|
// Why: temp paths on runners use 8.3 aliases (RUNNER~1). Git canonicalizes
|
|
// worktree paths, so a non-canonical HOME makes created worktrees invisible
|
|
// to the app's listing comparisons.
|
|
const isolatedHome = realpathSync.native(requestedIsolatedHome)
|
|
const app = await electron.launch({
|
|
executablePath: exePath,
|
|
args: [],
|
|
env: {
|
|
...cleanEnv,
|
|
// Packaged main honors ORCA_E2E_USER_DATA_DIR to relocate userData
|
|
// (logs/daemon/terminal-history) under a controlled dir.
|
|
...extraEnv,
|
|
ORCA_E2E_USER_DATA_DIR: userDataDir,
|
|
HOME: isolatedHome,
|
|
USERPROFILE: isolatedHome,
|
|
ORCA_E2E_HOME_DIR: isolatedHome,
|
|
ORCA_CODEX_SYSTEM_DEFAULT_REAL_HOME: '0'
|
|
}
|
|
})
|
|
// If firstWindow times out (the launched main never shows a window), the
|
|
// Electron process is still running — force-kill its tree before rethrowing so
|
|
// a driving failure never leaks an orphaned main to the CI job timeout.
|
|
let page
|
|
try {
|
|
page = await app.firstWindow({ timeout: 120_000 })
|
|
await page.waitForLoadState('domcontentloaded')
|
|
} catch (err) {
|
|
const pid = await resolveElectronMainPid(app)
|
|
if (pid) {
|
|
try {
|
|
execFileSync('taskkill', ['/pid', String(pid), '/T', '/F'], { stdio: 'ignore' })
|
|
} catch {
|
|
/* already gone */
|
|
}
|
|
}
|
|
throw err
|
|
}
|
|
return { app, page }
|
|
}
|
|
|
|
/** Resolve the packaged Electron main, optionally falling back to Playwright's child PID. */
|
|
export async function resolveElectronMainPid(
|
|
app,
|
|
{ allowLauncherFallback = true, timeoutMs = 5_000 } = {}
|
|
) {
|
|
let timeout
|
|
try {
|
|
// Why: packaged launchers can re-exec, leaving app.process() pointing at a
|
|
// dead stub while evaluate runs in the authoritative Electron main.
|
|
const pid = await Promise.race([
|
|
app.evaluate(() => process.pid),
|
|
new Promise((_, reject) => {
|
|
// Why: a wedged main connection is common on cleanup paths; resolving
|
|
// its authoritative PID must not consume the entire CI job timeout.
|
|
timeout = setTimeout(() => reject(new Error('main PID resolution timed out')), timeoutMs)
|
|
timeout.unref?.()
|
|
})
|
|
])
|
|
if (Number.isInteger(pid) && pid > 0) {
|
|
return pid
|
|
}
|
|
} catch {
|
|
/* the main connection may already be unavailable */
|
|
} finally {
|
|
clearTimeout(timeout)
|
|
}
|
|
// Why: crash proofs must fail closed rather than kill a packaged launcher stub
|
|
// and mistake its death for the authoritative Electron main crashing.
|
|
if (!allowLauncherFallback) {
|
|
return null
|
|
}
|
|
const fallbackPid = app.process()?.pid
|
|
return Number.isInteger(fallbackPid) && fallbackPid > 0 ? fallbackPid : null
|
|
}
|
|
|
|
/**
|
|
* Best-effort diagnostics dump when driving fails: a screenshot, the visible
|
|
* body text, and whether the e2e store is exposed (it is not in production
|
|
* builds). Written under `dir` so CI can upload it and reveal the actual
|
|
* post-launch DOM state. Never throws.
|
|
*/
|
|
export async function captureFailureDiagnostics(page, dir, label) {
|
|
const out = {}
|
|
try {
|
|
mkdirSync(dir, { recursive: true })
|
|
} catch {
|
|
return out
|
|
}
|
|
try {
|
|
await page.screenshot({
|
|
path: path.join(dir, `${label}.png`),
|
|
fullPage: false,
|
|
timeout: 10_000
|
|
})
|
|
out.screenshot = `${label}.png`
|
|
} catch {
|
|
/* renderer may be unresponsive */
|
|
}
|
|
try {
|
|
const info = await page.evaluate(() => ({
|
|
hasStore: typeof window.__store,
|
|
title: document.title,
|
|
url: location.href,
|
|
bodyText: (document.body?.innerText ?? '').slice(0, 4000),
|
|
testIds: Array.from(document.querySelectorAll('[data-testid]'))
|
|
.map((el) => el.getAttribute('data-testid'))
|
|
.filter((v, i, a) => v && a.indexOf(v) === i)
|
|
.slice(0, 80),
|
|
buttons: Array.from(document.querySelectorAll('button,[role="button"]'))
|
|
.map((el) => (el.getAttribute('aria-label') || el.textContent || '').trim())
|
|
.filter((v, i, a) => v && a.indexOf(v) === i)
|
|
.slice(0, 60),
|
|
tabs: Array.from(document.querySelectorAll('[data-testid="sortable-tab"]'))
|
|
.map((el) => ({
|
|
id: el.getAttribute('data-tab-id'),
|
|
title: el.getAttribute('data-tab-title'),
|
|
ariaLabel: el.getAttribute('aria-label'),
|
|
selected: el.getAttribute('aria-selected')
|
|
}))
|
|
.slice(0, 40),
|
|
activeElement: document.activeElement
|
|
? {
|
|
tag: document.activeElement.tagName,
|
|
className: document.activeElement.getAttribute('class'),
|
|
ariaLabel: document.activeElement.getAttribute('aria-label')
|
|
}
|
|
: null
|
|
}))
|
|
writeFileSync(path.join(dir, `${label}.json`), JSON.stringify(info, null, 2))
|
|
out.info = info
|
|
} catch {
|
|
/* renderer may be unresponsive */
|
|
}
|
|
return out
|
|
}
|
|
|
|
/** Wait until the visible terminal surface and its xterm container are mounted.
|
|
* An expected tab id prevents post-restore probes from accepting another tab. */
|
|
export async function waitForTerminalReady(page, timeoutMs = 60_000, terminalTabId = null) {
|
|
const selector = terminalTabId
|
|
? `[data-terminal-tab-id="${terminalTabId}"]:visible`
|
|
: TERMINAL_SURFACE_VISIBLE
|
|
const surface = page.locator(selector).first()
|
|
await surface.waitFor({ state: 'visible', timeout: timeoutMs })
|
|
await surface.locator(XTERM_CONTAINER_VISIBLE).first().waitFor({
|
|
state: 'visible',
|
|
timeout: timeoutMs
|
|
})
|
|
}
|
|
|
|
/**
|
|
* Get the app to an interactive terminal.
|
|
* - `allowCreate` true (first launch): if no terminal is visible, create a
|
|
* workspace from the seeded repo (or a new tab if a workspace already
|
|
* exists) — the drivable composer, not the native folder dialog.
|
|
* - `allowCreate` false (post-update relaunch): the session should be
|
|
* RESTORED, so only wait for the restored terminal — never create a second
|
|
* workspace (which would mask a broken restore).
|
|
*/
|
|
export async function ensureTerminal(page, { allowCreate = true, timeoutMs = 60_000 } = {}) {
|
|
const visibleTerminal = page.locator(TERMINAL_SURFACE_VISIBLE).first()
|
|
if (await visibleTerminal.isVisible().catch(() => false)) {
|
|
await waitForTerminalReady(page, timeoutMs)
|
|
return
|
|
}
|
|
if (!allowCreate) {
|
|
// Wait for the restored terminal to appear; a timeout here is a real
|
|
// (asserted) failure of session restore, not a driving gap.
|
|
await waitForTerminalReady(page, timeoutMs)
|
|
return
|
|
}
|
|
const newTab = page.getByRole(NEW_TAB_BUTTON.role, { name: NEW_TAB_BUTTON.name }).first()
|
|
if (await newTab.isVisible().catch(() => false)) {
|
|
await createTerminalTab(page)
|
|
return
|
|
}
|
|
await createWorkspaceFromSeededRepo(page, timeoutMs)
|
|
await waitForTerminalReady(page, timeoutMs)
|
|
}
|
|
|
|
/**
|
|
* Drive the "New workspace" composer to create a worktree from the single
|
|
* seeded project. The composer is in-app DOM (unlike the native Add-Project
|
|
* dialog): open it, choose the "Blank Terminal" mode so the worktree opens a
|
|
* plain terminal (not an agent), then submit "Create worktree".
|
|
*/
|
|
async function createWorkspaceFromSeededRepo(page, timeoutMs) {
|
|
await page
|
|
.getByRole(NEW_WORKSPACE_BUTTON.role, { name: NEW_WORKSPACE_BUTTON.name })
|
|
.first()
|
|
.click({ timeout: timeoutMs })
|
|
// Choose the plain-terminal mode (best-effort — if it is already the default
|
|
// or the label differs, the create below still produces a worktree).
|
|
await page
|
|
.getByRole('button', { name: 'Blank Terminal' })
|
|
.first()
|
|
.click({ timeout: 15_000 })
|
|
.catch(() => {})
|
|
// Submit. The create button's accessible name carries the shortcut hint
|
|
// ("Create worktreeCtrl"), so match by prefix; fall back to the documented
|
|
// Ctrl+Enter shortcut if the button is not directly clickable.
|
|
const created = await page
|
|
.getByRole('button', { name: /^Create worktree/ })
|
|
.last()
|
|
.click({ timeout: 15_000 })
|
|
.then(() => true)
|
|
.catch(() => false)
|
|
if (!created) {
|
|
await page.keyboard.press('Control+Enter')
|
|
}
|
|
}
|
|
|
|
const OVERLAY_DISMISS_LABELS = ['Got it', 'Dismiss setup scripts', 'Dismiss tip', 'Dismiss update']
|
|
|
|
/**
|
|
* Best-effort dismissal of the modals/banners that appear after creating a
|
|
* worktree (a full-screen "Got it" feature-tip modal, the setup-script prompt,
|
|
* update banner) and intercept all input over the terminal. Loops because tips
|
|
* can appear in sequence. Never throws.
|
|
*/
|
|
export async function dismissOverlays(page, rounds = 3) {
|
|
for (let i = 0; i < rounds; i++) {
|
|
let acted = false
|
|
for (const name of OVERLAY_DISMISS_LABELS) {
|
|
const btn = page.getByRole('button', { name }).first()
|
|
if (await btn.isVisible().catch(() => false)) {
|
|
await btn.click({ timeout: 3_000 }).catch(() => {})
|
|
acted = true
|
|
}
|
|
}
|
|
await page.keyboard.press('Escape').catch(() => {})
|
|
if (!acted) {
|
|
return
|
|
}
|
|
await page.waitForTimeout(400)
|
|
}
|
|
}
|
|
|
|
/** Create a new terminal tab via the New tab menu. Returns the count after. */
|
|
export async function createTerminalTab(page) {
|
|
await dismissOverlays(page, 1)
|
|
const before = await page.locator(SORTABLE_TAB).count()
|
|
await page
|
|
.getByRole(NEW_TAB_BUTTON.role, { name: NEW_TAB_BUTTON.name })
|
|
.first()
|
|
.click({ force: true })
|
|
await page.getByRole('menuitem', { name: NEW_TERMINAL_ITEM }).first().click({ force: true })
|
|
await page.waitForFunction(
|
|
({ selector, prev }) => document.querySelectorAll(selector).length > prev,
|
|
{ selector: SORTABLE_TAB, prev: before },
|
|
{ timeout: 10_000 }
|
|
)
|
|
await waitForTerminalReady(page)
|
|
return page.locator(SORTABLE_TAB).count()
|
|
}
|
|
|
|
/** Cheap session identifiers: the rendered tab ids. */
|
|
export async function listTabIds(page) {
|
|
return page
|
|
.locator(SORTABLE_TAB)
|
|
.evaluateAll((tabs) =>
|
|
tabs.map((t) => t.getAttribute('data-tab-id')).filter((id) => Boolean(id))
|
|
)
|
|
}
|
|
|
|
/**
|
|
* Focus the live terminal so keystrokes reach the shell. Clicking the visible
|
|
* xterm surface is what actually gives xterm keyboard focus — focusing the
|
|
* off-screen helper textarea alone does not, which is why typed input was being
|
|
* dropped. Click the pane, then focus the helper textarea as a belt-and-braces.
|
|
*/
|
|
export async function focusActiveTerminal(page, terminalTabId = null) {
|
|
// A feature-tip modal can appear late and swallow keystrokes; clear any before
|
|
// focusing so typed commands actually reach the shell.
|
|
for (const name of OVERLAY_DISMISS_LABELS) {
|
|
const btn = page.getByRole('button', { name }).first()
|
|
if (await btn.isVisible().catch(() => false)) {
|
|
await btn.click({ timeout: 2_000 }).catch(() => {})
|
|
}
|
|
}
|
|
const selector = terminalTabId
|
|
? `[data-terminal-tab-id="${terminalTabId}"]:visible`
|
|
: TERMINAL_SURFACE_VISIBLE
|
|
const surface = page.locator(selector).first()
|
|
const click = surface.click({ position: { x: 24, y: 24 }, timeout: 15_000 })
|
|
// Why: an exact-tab proof must fail closed if that restored surface vanishes;
|
|
// typing into whichever element retained focus could falsely target another tab.
|
|
await (terminalTabId ? click : click.catch(() => {}))
|
|
// Scope the helper textarea to the visible surface so focus can't land on a
|
|
// hidden duplicate pane's textarea (which would silently swallow keystrokes).
|
|
const input = surface.locator(XTERM_INPUT).last()
|
|
const focus = input.focus()
|
|
await (terminalTabId ? focus : focus.catch(() => {}))
|
|
return input
|
|
}
|
|
|
|
/** Type a line and submit it (Enter → \r submits in the shell). */
|
|
export async function typeLine(page, text, terminalTabId = null) {
|
|
await focusActiveTerminal(page, terminalTabId)
|
|
await page.keyboard.type(text)
|
|
await page.keyboard.press('Enter')
|
|
}
|
|
|
|
/** Send Ctrl+C to the active terminal. */
|
|
export async function sendCtrlC(page, terminalTabId = null) {
|
|
await focusActiveTerminal(page, terminalTabId)
|
|
await page.keyboard.press('Control+C')
|
|
}
|
|
|
|
/**
|
|
* Run a PowerShell command inside the active terminal by invoking a nested
|
|
* powershell.exe. The command is wrapped in double quotes for the OUTER
|
|
* interactive shell (also pwsh), which would otherwise expand `$var`, `$(...)`
|
|
* and consume backticks before the nested shell sees them — so escape backticks,
|
|
* quotes, and `$`. Without the `$` escape, `while($true)` reaches the nested
|
|
* shell as `while(True)` and never runs (the bug that silently broke every
|
|
* loop/heartbeat probe while simple `$`-free commands worked).
|
|
*/
|
|
export async function runShellCommand(page, psCommand) {
|
|
const escaped = psCommand.replace(/`/g, '``').replace(/"/g, '`"').replace(/\$/g, '`$')
|
|
await typeLine(page, `powershell.exe -NoProfile -NonInteractive -Command "${escaped}"`)
|
|
}
|
|
|
|
/**
|
|
* Start the long-running marker in the active terminal: sets the canary window
|
|
* title, records its own PID, and heartbeats a file every 500ms. Returns the
|
|
* command string (the caller reads the pid file to learn the marker PID).
|
|
*/
|
|
export async function startMarker(page, { canary, pidFile, heartbeatFile }) {
|
|
const script = [
|
|
`$host.UI.RawUI.WindowTitle='${canary}'`,
|
|
`Set-Content -LiteralPath '${pidFile}' -Value $PID`,
|
|
`while($true){ [System.IO.File]::WriteAllText('${heartbeatFile}',(Get-Date).ToString('o')); Start-Sleep -Milliseconds 500 }`
|
|
].join('; ')
|
|
await runShellCommand(page, script)
|
|
return script
|
|
}
|
|
|
|
/**
|
|
* Best-effort read of the active terminal's visible text, for the cold-restore
|
|
* scrollback fidelity check. Prefers the SerializeAddon when the build happens
|
|
* to expose paneManagers; falls back to DOM rows (populated only under the DOM
|
|
* renderer, so this may be empty under WebGL — hence best-effort).
|
|
*/
|
|
export async function readTerminalTextBestEffort(page) {
|
|
return page.evaluate(() => {
|
|
const managers = window.__paneManagers
|
|
if (managers && typeof managers.forEach === 'function') {
|
|
let out = ''
|
|
managers.forEach((m) => {
|
|
const pane = m.getActivePane?.() ?? m.getPanes?.()[0]
|
|
const text = pane?.serializeAddon?.serialize?.()
|
|
if (text) {
|
|
out += text
|
|
}
|
|
})
|
|
if (out) {
|
|
return out
|
|
}
|
|
}
|
|
return Array.from(document.querySelectorAll('.xterm-rows'))
|
|
.map((el) => el.textContent ?? '')
|
|
.join('\n')
|
|
})
|
|
}
|
|
|
|
/**
|
|
* Close the app gracefully; force-kill its process tree on timeout. Mirrors
|
|
* tests/e2e/helpers/electron-process-shutdown.ts so the daemon (detached) is
|
|
* left alive exactly as a normal quit would.
|
|
*/
|
|
export async function closeApp(app, timeoutMs = 10_000) {
|
|
// A partially-created session (launch failed before assignment) passes undefined.
|
|
if (!app) {
|
|
return
|
|
}
|
|
const mainPid = await resolveElectronMainPid(app)
|
|
let closeTimeout
|
|
try {
|
|
await Promise.race([
|
|
app.close(),
|
|
new Promise((_, reject) => {
|
|
closeTimeout = setTimeout(() => reject(new Error('close timeout')), timeoutMs)
|
|
closeTimeout.unref?.()
|
|
})
|
|
])
|
|
} catch {
|
|
if (mainPid) {
|
|
try {
|
|
execFileSync('taskkill', ['/pid', String(mainPid), '/T', '/F'], { stdio: 'ignore' })
|
|
} catch {
|
|
/* already gone */
|
|
}
|
|
}
|
|
} finally {
|
|
// Why: successful closes must not retain a timeout closure or keep a shared
|
|
// harness process alive until the failure deadline expires.
|
|
clearTimeout(closeTimeout)
|
|
}
|
|
}
|