Commit Graph
6 Commits
Author SHA1 Message Date
Neil a781a602a8 test: retire duplicate cases that replay an owner across a re-export or provider shim (#24114)
Resolves 208 candidate pairs where the same case title appears verbatim in two or more
files, produced by a repo-wide scan calibrated against a known positive. 46 case
declarations removed across 32 files, 798 lines gone. No file deleted whole, no
production code touched.

The headline result is the measurement, not the deletions: across the three buckets that
reported in detail, the signal ran roughly 86% false-positive (3/42, 9/42, and the rest).
It has good recall and poor precision, and it reorders a reading queue rather than
replacing one. Calibrating a detector against a known positive proves recall, not
precision.

What the deletions were:

- Duplicate invocation through a re-export shim. `native-chat-tool-summary.ts` is a
  ten-line `export {...} from '../../../../shared/native-chat-tool-summary'`, and
  `agent-status.ts:161` is `export { isExplicitAgentStatusFresh } from
  './pane-agent-evidence'`. Cases on the shim side were byte-equivalent to the owner's
  with no rendering or transport hop.
- Provider-local replays of a shared helper: three `repository-ref` providers that are
  each `createRemoteRefProbeCache(parseXRef)` and contribute nothing to transient
  handling; two `local-pty` and `daemon/session` tables replaying
  `shell-startup-output-scanner`, whose owner additionally checks every split point.
- A reader-side replay of store policy. `runtime-worktree-agent-rows-structured.test.ts`
  asserted an attention-to-blocked mapping; the reader contains zero `attention` or
  `blocked` tokens and copies `state` through. The mapping lives in
  `structuredAgentSessionAgentStatus`. Consistent with
  `docs/reference/agent-status-store.md`: readers keep only presentation policy.
- Constructor-only subclass duplication: the shared capability-cache case is covered by
  `codex-app-server-capability-cache.test.ts`, whose ten cases include the identical
  title plus all four risks `docs/reference/git-compatibility.md` names — first fallback,
  later cached call, concurrent probes, per-host isolation.
- A private predicate duplicated at a real boundary, varying only a path passed straight
  into the shared predicate.

Why most pairs were KEPT, because the false positives are principled rather than noise:

- Two independent execution hosts. `src/relay/git-handler-*` and `src/main/git/*` are
  separate Git implementations that cannot import each other and hold separate capability
  caches, exactly as the compatibility doc requires; the repo already ships
  `status-branch-line-total-relay-parity.test.ts` to pin the duality deliberately. Neither
  side's argv, timeout or cache regression is visible to the other.
- Deliberately duplicated production siblings: Codex vs Claude (different account fields,
  different CLIs, different wire protocols), gitea vs bitbucket (`/pulls/42` vs
  `/pullrequests/42`), gl-utils vs gh-utils (separate in-flight maps). Same contract
  shape, different implementations — an identical title is the correct naming.
- Shared-predicate consumers: one side tests the predicate, the other tests a caller's
  wiring to it. A caller that forgot to call the predicate passes the shared test.

In a codebase with intentional provider and host symmetry, identical test titles are
expected, and the signal cannot distinguish "copied" from "parallel by design" because
both produce the same prose. Only reading both bodies separates them.

Verified: 6,968 desktop test files pass; the three modified mobile files pass (39 cases);
`check-reliability-gates.mjs` 140 gates; nothing under
`mobile/src/test-support/rpc-recording/` or `mobile/rpc-foundation/goldens/` touched.

62 local failures across 12 files were each accounted for and none is caused by this
change: `browser-manager-tab-identity`, `browser-manager-viewport-ownership`,
`session-scanner-codex-workers` and `managed-hook-script-refresh` all fail identically in
a pristine `origin/main` worktree; five `mobile-web-app-*-render` tests need Playwright
browsers this machine lacks; `structured-agent-session-restart-ownership` and
`ssh-remote-commands` pass in isolation and fail only under concurrent load.
2026-09-30 03:30:23 -07:00
OrcaWinandm4air 8416e8de10 refactor(persistence): retire ordinary JSON profile writes (#23202)
* refactor(persistence): retire ordinary JSON profile writes

Require SQLite for writable profiles and keep import, compatibility export, and recovery in a documented legacy-json boundary.

* fix(cli): preserve dynamic profile imports in release output

* test(persistence): exercise SQL races and verify packaged CLI imports

* test(persistence): consolidate shared fixture imports

* test(persistence): close SQLite fixtures before cleanup and await launcher output

* test(automations): use SQLite fixtures for dispatch fencing and skip coalescing

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-09-26 12:56:32 -07:00
OrcaWin 82412dab8b Persist profile state in SQLite with background writes (#22612)
Migrate profile state to SQLite and move writes and backups into a background worker. Acknowledge terminal, SSH and automation changes only after durable saves. Preserve JSON import, recovery, rollback and compatibility exports.

Validate migration, worker failures, maintenance, cross-profile moves and terminal lifetime races with unit, integration and end-to-end coverage.
2026-09-25 22:47:33 -07:00
Neil a2aea5d0b0 fix(persistence): sweep rows owned by deregistered repo ids at load
Deregistering a project stranded every row it owned. Each pruning path is
gated on the repo still being in `state.repos`, so once an id leaves the
catalogue its metadata, identity aliases, lineage and session rows became
unreachable forever -- and on a paired client they rendered as phantom
worktrees under an "Unknown" project.

Reconcile against the repo catalogue on load instead: any repo id that owns
rows but is absent from `state.repos` has its rows removed through the same
path `removeProject` uses. Host-independent and session-independent, because
an orphan has no owner that could object -- which is also why this reaches a
client's mirror of a remote host's session partition, something no local
removal can do.

Only a full `<repoId>::<path>` locator seeds the orphan set; bare keys can be
folder workspace ids or repo-keyed revisions, and guessing wrong there would
delete live state. `retiredWorktreeNamesByRepo` is deliberately untouched so a
re-added repo cannot reissue a name onto a cwd that still holds a prior
occupant's agent state.

Test fixtures that wrote worktree rows without registering their repo were
relying on orphans surviving a reload; they now register the repo they name.

Refs #17776
2026-09-01 17:20:17 -07:00
Neil 7ae916cebd perf(worktrees): batch-prune stale local metadata (#17278) 2026-08-29 16:06:15 -07:00
Neil 9367169888 refactor(tests): split every oversized test file off the max-lines suppression list (#14728)
* refactor(tests): split oversized test files off the max-lines suppression list

Every `*.test.ts`/`*.spec.ts` that carried an `eslint/oxlint-disable max-lines`
directive is now split into focused, behavior-scoped suites that fit the 800-line
test budget, with shared setup extracted into co-located `*-test-harness.ts` /
`*-test-fixtures.ts` modules (300-line budget). 83 files became ~930; the largest
output is 797 effective lines. `orca-runtime.test.ts` is intentionally untouched.

Test bodies were moved by scripted line-range slicing rather than retyped, so
assertions are byte-identical. The only permitted body edits were mechanical
rebinding where a shared value moved into a harness (e.g. `tmpHome` ->
`homes.tmpHome`).

Registries that enumerate test files were updated in lockstep:
- config/max-lines-baseline.txt: pruned 341 -> 258 entries (all 83 removed).
- config/reliability-gates.jsonc: 33 gates repointed at the split files, with
  assertionRefs split per file where a gate's coverage now spans several.
- .github/workflows/pr.yml: the real-zsh lane now lists the 4 split files that
  actually exercise zsh, so they keep running in the dedicated shell lane.

Also renamed agent-hooks `server-test-fixtures.ts` to `server.test-fixtures.ts`
so the global-fetch call-site audit keeps skipping it, and added `.js` extensions
to the CLI suites' dynamic harness imports (node16 resolution) to unbreak
`build:cli`.

Verification: full suite 52,449 passing vs 52,448 at baseline with zero
assertions lost; `pnpm lint`, `pnpm typecheck`, and `pnpm build:cli` all exit 0;
the terminal-pane e2e spec runs 31/31 headless.

* refactor(tests): split hook-idle arbitration suite that oxfmt pushed over budget

The pre-commit oxfmt pass reflowed pty-connection-hook-idle-arbitration.test.ts
to 811 effective lines, 11 over the test budget. Split the hook-completion side
effect and replacement-agent veto cases into their own suite; both files now sit
well under the cap and the 15 tests are unchanged.

* test: port upstream test changes into the split files after rebase

Rebasing onto main surfaced 27 tests that main had added to files this branch
deleted, plus edits to tests that had already moved. Taking the deletion side of
those modify/delete conflicts would have dropped that coverage silently, so each
upstream change is ported into the split file that now owns the behavior — for
example main's six orchestration mailbox tests land across orchestration-runs,
-send, and -check.

Also repoints `orchestration.notification-mailbox-consistency`, a gate main added
after this branch's gate remap, at those same three split files, and re-prunes
the max-lines baseline against main's (257 entries).

Verified: all 27 upstream test titles present; full suite 52,761 passing with the
only diff vs baseline being 12 tests main itself removed and 3 that moved from
skipped to passing; lint and typecheck exit 0.

* fix(test): flush pending continuations before tearing down terminal test globals

CI shard 5/16 failed on both Node 24 and 26 with `ReferenceError: window is not
defined` from pty-connection.ts, surfacing through
pty-connection-daemon-snapshot-replay.test.ts.

The reattach/settle chains `await` a real promise and then touch `window.api`.
Under fake timers those continuations cannot run, so they only become schedulable
once restoreTerminalTestGlobals() switches back to real timers — which previously
happened immediately before `delete globalThis.window`, so a late continuation
threw and failed the whole file. Flush async ticks in that window instead.

This is latent in the source rather than new: the pre-split 25k-line file kept
running other tests after these, which gave the chains time to settle before
teardown. Splitting the file moved teardown directly behind them.

* fix(test): keep an inert window after terminal test teardown instead of deleting it

The async-tick flush was not enough: the reattach/settle chain can resolve after
teardown regardless of how long we drain, so CI shard 5/16 still failed with
`ReferenceError: window is not defined` from pty-connection.ts.

A real renderer never loses `window`, so deleting it was the artificial part.
Swap in an inert proxy whose properties resolve to callables and whose calls
resolve to undefined, making a late `window.api.pty.*` call a harmless no-op.
The next test replaces it wholesale via installTerminalTestGlobals(), and no test
asserts that `window` is absent.
2026-08-15 00:54:20 -07:00