mirror of
https://github.com/stablyai/orca.git
synced 2026-09-22 16:02:32 +00:00
* test(e2e): make Codex typing-latency harness measure real echo latency The local Codex typing-latency spec produced meaningless numbers. Four defects, all fixed here: 1. False-positive readiness. `/Ask Codex|OpenAI/i` matched "OpenAI's command-line coding agent" on the *sign-in* screen, so the test went "ready" against a login prompt and measured typing into a non-composer. Now gated on the composer status bar (`/Context \d+% used/i`), which only the live composer draws. Banner text is unusable: the serialized buffer interleaves ANSI escapes through those glyphs. 2. Missing auth. The E2E profile runs an isolated HOME with a managed CODEX_HOME that has no auth.json, guaranteeing the sign-in screen. The launch now pins the real ~/.codex, and skips with a clear message when auth.json is absent instead of silently measuring a login screen. 3. Measurement overhead swamped the signal. Per-key latency was measured by polling getTerminalContent() every 5ms, so each sample was real echo latency + full buffer serialize + CDP round-trip + poll granularity. Measurement now happens entirely in-renderer: an in-page hook stamps performance.now() on keydown (window capture phase, before xterm forwards to the PTY) and again in xterm's onWriteParsed once the glyph is in the viewport, with onRender giving a separate time-to-paint. Samples are drained in one page.evaluate after typing ends — zero CDP round-trips inside the measured window. 4. Thresholds were meaningless (median<150ms / worst<500ms). Replaced with p50<35 / p95<60 / max<120, based on 10 local runs. Also: 60 keystrokes instead of 24 with the first 10 discarded as warmup, p50/p95/max instead of a lone median, lowercase-only input so the slash and file-mention popups can't perturb later keys, an assertion that no keystroke went unechoed, and a terminal dump on readiness failure. Measured (10 local runs, headless, real Codex 0.145.0): echo (key->parse) p50 21.6-22.6ms, p95 23.2-41.5ms, max 23.4-58.7ms paint (key->render) p50 25.5-32.9ms, p95 34.3-49.7ms A plain-shell control on the same probe reads p50 2.0ms / p95 3.0ms, confirming the ~22ms is Codex composer redraw cost rather than a harness floor — the old harness reported ~29-30ms for everything. Co-authored-by: Orca <help@stably.ai> * test(e2e): widen Codex latency tail budgets and assert terminal focus Follow-up calibration over ~20 local runs: the per-key distribution is unimodal at p50 21.3-22.7ms with rare isolated spikes to ~90-125ms that are not a steady-state shift. Tail budgets move to p95<80 / max<150 so only a sustained regression fails; p50<35 still gates the steady state. Also assert the xterm helper textarea actually took focus. One run typed all 60 keys with only 5 parse events because focus was lost, which previously surfaced as an opaque sample-count mismatch. Co-authored-by: Orca <help@stably.ai> --------- Co-authored-by: Orca <help@stably.ai>