mirror of
https://github.com/stablyai/orca.git
synced 2026-10-04 08:02:09 +00:00
Resolves 208 candidate pairs where the same case title appears verbatim in two or more
files, produced by a repo-wide scan calibrated against a known positive. 46 case
declarations removed across 32 files, 798 lines gone. No file deleted whole, no
production code touched.
The headline result is the measurement, not the deletions: across the three buckets that
reported in detail, the signal ran roughly 86% false-positive (3/42, 9/42, and the rest).
It has good recall and poor precision, and it reorders a reading queue rather than
replacing one. Calibrating a detector against a known positive proves recall, not
precision.
What the deletions were:
- Duplicate invocation through a re-export shim. `native-chat-tool-summary.ts` is a
ten-line `export {...} from '../../../../shared/native-chat-tool-summary'`, and
`agent-status.ts:161` is `export { isExplicitAgentStatusFresh } from
'./pane-agent-evidence'`. Cases on the shim side were byte-equivalent to the owner's
with no rendering or transport hop.
- Provider-local replays of a shared helper: three `repository-ref` providers that are
each `createRemoteRefProbeCache(parseXRef)` and contribute nothing to transient
handling; two `local-pty` and `daemon/session` tables replaying
`shell-startup-output-scanner`, whose owner additionally checks every split point.
- A reader-side replay of store policy. `runtime-worktree-agent-rows-structured.test.ts`
asserted an attention-to-blocked mapping; the reader contains zero `attention` or
`blocked` tokens and copies `state` through. The mapping lives in
`structuredAgentSessionAgentStatus`. Consistent with
`docs/reference/agent-status-store.md`: readers keep only presentation policy.
- Constructor-only subclass duplication: the shared capability-cache case is covered by
`codex-app-server-capability-cache.test.ts`, whose ten cases include the identical
title plus all four risks `docs/reference/git-compatibility.md` names — first fallback,
later cached call, concurrent probes, per-host isolation.
- A private predicate duplicated at a real boundary, varying only a path passed straight
into the shared predicate.
Why most pairs were KEPT, because the false positives are principled rather than noise:
- Two independent execution hosts. `src/relay/git-handler-*` and `src/main/git/*` are
separate Git implementations that cannot import each other and hold separate capability
caches, exactly as the compatibility doc requires; the repo already ships
`status-branch-line-total-relay-parity.test.ts` to pin the duality deliberately. Neither
side's argv, timeout or cache regression is visible to the other.
- Deliberately duplicated production siblings: Codex vs Claude (different account fields,
different CLIs, different wire protocols), gitea vs bitbucket (`/pulls/42` vs
`/pullrequests/42`), gl-utils vs gh-utils (separate in-flight maps). Same contract
shape, different implementations — an identical title is the correct naming.
- Shared-predicate consumers: one side tests the predicate, the other tests a caller's
wiring to it. A caller that forgot to call the predicate passes the shared test.
In a codebase with intentional provider and host symmetry, identical test titles are
expected, and the signal cannot distinguish "copied" from "parallel by design" because
both produce the same prose. Only reading both bodies separates them.
Verified: 6,968 desktop test files pass; the three modified mobile files pass (39 cases);
`check-reliability-gates.mjs` 140 gates; nothing under
`mobile/src/test-support/rpc-recording/` or `mobile/rpc-foundation/goldens/` touched.
62 local failures across 12 files were each accounted for and none is caused by this
change: `browser-manager-tab-identity`, `browser-manager-viewport-ownership`,
`session-scanner-codex-workers` and `managed-hook-script-refresh` all fail identically in
a pristine `origin/main` worktree; five `mobile-web-app-*-render` tests need Playwright
browsers this machine lacks; `structured-agent-session-restart-ownership` and
`ssh-remote-commands` pass in isolation and fail only under concurrent load.
107 lines
4.5 KiB
JavaScript
107 lines
4.5 KiB
JavaScript
import { describe, expect, it } from 'vitest'
|
|
import {
|
|
createDailyBuildVersion,
|
|
formatDailyReleaseName,
|
|
nextDailyBuildNumber
|
|
} from './daily-build-version.mjs'
|
|
import { compareAppVersions } from '../../src/shared/app-version'
|
|
|
|
describe('createDailyBuildVersion', () => {
|
|
it('stamps the version with a zero-padded UTC timestamp', () => {
|
|
expect(createDailyBuildVersion('1.4.160', new Date('2026-07-28T13:00:00Z'))).toBe(
|
|
'1.4.160-daily.202607281300'
|
|
)
|
|
})
|
|
|
|
// Why: main's package.json carries the in-flight RC tail. Keeping it would make
|
|
// every daily semver-NEWER than the RC it was cut from (1.4.160-rc.3-daily.X >
|
|
// 1.4.160-rc.3), so an ordinary RC-channel check would offer untested daily
|
|
// builds to RC users. Dropping it parks dailies below both rc.N and stable,
|
|
// reachable only by an explicit pinned jump.
|
|
it('drops an in-flight rc tail so dailies never outrank the rc series', () => {
|
|
const version = createDailyBuildVersion('1.4.160-rc.3', new Date('2026-07-28T13:00:00Z'))
|
|
expect(version).toBe('1.4.160-daily.202607281300')
|
|
expect(compareAppVersions(version, '1.4.160-rc.3')).toBeLessThan(0)
|
|
expect(compareAppVersions('1.4.160-rc.3-daily.202607281300', '1.4.160-rc.3')).toBeGreaterThan(0)
|
|
})
|
|
|
|
// Why: 'daily' sorts before 'hourly' alphabetically, so a daily of the same
|
|
// base never outranks an hourly — both stay below rc/stable and only the
|
|
// channel picker offers them.
|
|
it('sorts below the hourly build of the same base version', () => {
|
|
expect(
|
|
compareAppVersions('1.4.160-daily.202607281300', '1.4.160-hourly.202607281400')
|
|
).toBeLessThan(0)
|
|
})
|
|
|
|
it('rejects invalid input', () => {
|
|
expect(() => createDailyBuildVersion('nope', new Date())).toThrow(/valid semver/)
|
|
expect(() => createDailyBuildVersion('1.4.160', new Date('nope'))).toThrow(/invalid/)
|
|
})
|
|
})
|
|
|
|
describe('formatDailyReleaseName', () => {
|
|
const name = (iso, buildNumber = 1, commit = 'e698241abcde') =>
|
|
formatDailyReleaseName('1.4.163-daily.x', buildNumber, commit, new Date(iso))
|
|
|
|
it('renders version, number, Pacific timestamp, and short sha', () => {
|
|
// 17:15 UTC is 10:15AM PDT in July (the daily cut time).
|
|
expect(name('2026-07-28T17:15:00Z')).toBe('1.4.163 • 01 • Jul 28, 10:15AM • e698241')
|
|
})
|
|
|
|
// Why both sides of DST: the tag's stamp is UTC and the title is Pacific, so
|
|
// the offset between them is not a constant. A test pinned to one season would
|
|
// pass all summer and start failing in November.
|
|
it('follows the Pacific offset across DST', () => {
|
|
expect(name('2026-01-15T18:15:00Z')).toBe('1.4.163 • 01 • Jan 15, 10:15AM • e698241')
|
|
expect(name('2026-07-28T17:15:00Z')).toBe('1.4.163 • 01 • Jul 28, 10:15AM • e698241')
|
|
})
|
|
|
|
it('rejects a build number that is not a positive integer', () => {
|
|
expect(() => name('2026-07-28T17:15:00Z', 0)).toThrow(/positive integer/)
|
|
expect(() => name('2026-07-28T17:15:00Z', -1)).toThrow(/positive integer/)
|
|
expect(() => name('2026-07-28T17:15:00Z', 1.5)).toThrow(/positive integer/)
|
|
})
|
|
|
|
it('rejects an invalid timestamp', () => {
|
|
expect(() => formatDailyReleaseName('1.4.163', 1, 'abcdefg', new Date('nope'))).toThrow(
|
|
/invalid/
|
|
)
|
|
})
|
|
})
|
|
|
|
describe('nextDailyBuildNumber', () => {
|
|
const titles = [
|
|
'1.4.163 • 01 • Jul 28, 6:00AM • e698241',
|
|
'1.4.163 • 02 • Jul 29, 6:00AM • aaaaaaa',
|
|
'1.4.163 • 09 • Aug 01, 6:00AM • bbbbbbb'
|
|
]
|
|
|
|
it('continues the series for the version being built', () => {
|
|
expect(nextDailyBuildNumber('1.4.163', titles)).toBe(10)
|
|
})
|
|
|
|
it('restarts at 1 when the base version moves', () => {
|
|
expect(nextDailyBuildNumber('1.4.164', titles)).toBe(1)
|
|
expect(
|
|
nextDailyBuildNumber('1.4.164', [...titles, '1.4.164 • 01 • Aug 02, 6:00AM • ccccccc'])
|
|
).toBe(2)
|
|
})
|
|
|
|
// Why max and not count: pruning trims to DAILY_RETAIN_COUNT, so counting
|
|
// would roll backwards and reissue a number already used.
|
|
it('takes the highest number, not the count', () => {
|
|
expect(nextDailyBuildNumber('1.4.163', ['1.4.163 • 09 • Jul 31, 6:00AM • e698241'])).toBe(10)
|
|
})
|
|
|
|
it('starts at 1 with no history at all', () => {
|
|
expect(nextDailyBuildNumber('1.4.163')).toBe(1)
|
|
expect(nextDailyBuildNumber('1.4.163', [])).toBe(1)
|
|
})
|
|
|
|
it('ignores titles that are not this version', () => {
|
|
expect(nextDailyBuildNumber('1.4.16', ['1.4.163 • 09 • Aug 01, 6:00AM • bbbbbbb'])).toBe(1)
|
|
expect(nextDailyBuildNumber('1.4.163', ['v1.4.163-daily.202607311300', null, ''])).toBe(1)
|
|
})
|
|
})
|