mirror of
https://github.com/stablyai/orca.git
synced 2026-10-01 00:02:10 +00:00
Completes the `src/main` audit. Sweep over `src/main/runtime` (flat, rpc,
orchestration, relay, push) and flat `src/main`. 42 case declarations removed
across 24 files, 783 lines gone. No file deleted whole, no production code touched.
Highest-yield area by far was `runtime/orchestration` (21 cases from 151 files);
the rest of runtime measured under 1%.
What went, by pattern:
- Tests whose subject is the test itself: a source-scanning boundary test with no
production import at all, three of whose cases checked its own regex against
strings it declares; and a benchmark whose own simulation contains the
short-circuit it asserts, with `expect(wouldHaveBeen).toBe(60000)` comparing the
test's own arithmetic.
- Identity copiers: seven db cases where every asserted value is the literal input
(`type: 'question'` in, `type === 'question'` out).
- Permanently skipped tests for behavior that does not exist — two cases carrying
`// TODO: inline restore on re-subscribe not yet implemented`. A skipped test for
an unimplemented feature can never fail; it is a note in test syntax.
- Duplicate invocations of a contract owned at a stronger boundary, including three
reset scopes and two dependency-promotion cases owned by dedicated suites.
- Registration manifests whose every entry is referenced by other tests.
- Table rows and cases varying a field production never reads: a mobile tab-restore
case varying `clientCapabilities`, which the mobile branch does not consult, and
a Windows worker case in a module with no platform input at all.
- Names promising more than the input exercises: a case titled for forged AppImage
variables whose body only removes `AppRun`, byte-identical to a row of the
`it.each` table twenty lines above.
- Negative controls passing for an unrelated reason: a foreign-pane rejection whose
fixture also differs in sender handle, so the handle guard can reject it.
- Self-comparisons, including `format(m, { authority: 'current' }) === format(m)`.
Two deletions were justified by the wrong argument and kept only after checking a
better one. "A generated-catalog check gates registration" is false: that script
prevents the catalog and dispatcher from drifting apart, so a developer who removes
a method regenerates the catalog and the check passes. Like a type annotation over
an interface and its implementing class, it verifies internal consistency and
cannot see a declaration and its use removed together. Both inventories go on
per-entry evidence instead — every name is referenced by other tests.
The same rule kept a 37-entry terminal-method inventory in the same wave, because
some of its entries are pinned solely by it. An inventory is a ratchet if and only
if at least one entry is pinned solely by it; that is evidence per entry, not a
verdict by shape.
Kept deliberately: everything a gate cites, checked by case title and not only by
file path; source-scanning ratchets that pair their negative assertion with a
positive one against real source (the surviving boundary case asserts the pattern
still matches the writer module, so a silently-broken regex goes red); a
destructive-delete PTY waiver guard; `it.fails` markers, which go red if the bug is
fixed; and `it.skipIf(platform)` cases, which do run on other hosts.
Coverage is partial and stated as such: 1,195 files in scope, roughly 660 read
case-by-case, the remainder title-scanned and mechanically triaged. Unread paths are
recorded for a later sweep rather than assumed clean.
Verified: `check-reliability-gates.mjs` (140 gates), `check:code-quality:changed`
(0 new findings). No `pnpm tc` needed since no production file changed. A local run
of `src/main` surfaced 42 failures in three files this change does not touch
(`browser-manager-tab-identity`, `browser-manager-viewport-ownership`,
`session-scanner-codex-workers`); all 42 reproduce on a pristine `origin/main`
worktree, so they predate this wave and are environment-dependent locally — CI was
green on main at `2d85fdc753e2`.