mirror of
https://github.com/stablyai/orca.git
synced 2026-09-29 16:02:50 +00:00
* fix(worktrees): keep retirement tombstones across project and SSH target re-add Generated workspace names are retired so a name is never reissued onto a cwd that still holds another workspace's Claude/Codex history. Two re-add paths lost that record. STA-4449 (local): retirement was stored only under `repo.id`. Removing a project deletes that row and re-adding the same path mints a new id, so the new repo starts with an empty registry. The on-disk backfill normally re-seeds local repos, but it cannot recover a name whose only surviving evidence is a Codex rollout JSONL — those are deliberately not scanned — so a name spent under Codex with its workspace directory gone came back. STA-4491 (SSH): the second, path-derived copy embedded the SSH target row id. Row ids are minted fresh on every re-add, so `ssh:ssh-old:...` became `ssh:ssh-new:...` and `reassignSshTargetId` migrated other carrier state but not the retirement namespaces. Keying the store on the namespace instead of `repo.id` was rejected in `nestWorkspaces`, `worktreeBasePath` and `repo.path`, so a settings toggle would orphan every retirement at once. `repo.id` therefore stays primary and the path-derived namespace stays a mirror — a settings toggle loses the mirror but keeps the repo row, a re-add loses the repo row but keeps the mirror, and reads union both. - Mirror local repos into the namespace too, not just remote ones. - Key the namespace's host half on the SSH endpoint (host+port+username), the thing that actually decides which filesystem a path lands on, instead of the target row id. Reads also accept the pre-identity key so an upgrade keeps tombstones it already wrote, and `reassignSshTargetId` re-keys the rest. - Cap the namespace map, which by design outlives the repos that wrote it and so has nothing to prune it per repo. Endpoint identity is extracted from ssh-target-readoption.ts, which already compared these fields for exactly the same reason, so re-adoption and retirement cannot drift apart. * fix(worktrees): copy shared SSH endpoint retirements instead of moving them An endpoint identity is not owned by the target row that rotates: nothing dedupes SSH targets by host|port|username, so a second live target can still resolve to the same host. Moving the bucket stripped that target's tombstones and reissued a path whose agent history is still on disk. Row-id identities stay a move — reassignment leaves nothing pointing at them. * fix(worktrees): carry retirement mirror across in-place SSH endpoint edits Config sync rewrites host/port/username on the existing target row and a runtime-owned target takes a fresh address from every provision, both keeping the row id. No re-adoption runs, so nothing carried the endpoint-keyed mirror across and it stranded — strictly worse than the pre-change key, which was the row id and was invariant under these edits. Also bound the map after a migration: a retained source bucket grows it, so the cap has to be applied there too, and compare registries by membership rather than size so an uncompacted destination cannot trade a folded name for a new one and read as unchanged. * fix(worktrees): skip retirement migration for on-demand runtime targets An on-demand VM is discarded between provisions, so its fresh address reaches an empty filesystem and a reissued name collides with nothing. Migrating there would spend names against history that no longer exists, and because each provision mints another address it would add a namespace bucket per run, evicting the real tombstones of local and ordinary SSH repos. * fix(worktrees): stop the namespace cap evicting what a migration just wrote Two defects with one root cause. Assigning to an existing key leaves it in its original insertion slot, so a merged destination kept the oldest position and the trim deleted the bucket it had just enriched. Retained source buckets are older than the destinations a copy appends, so at the cap the trim removed exactly the sources the copy existed to keep — silently turning it back into a move. The trim now exempts the keys the migration wrote or deliberately kept. Also stop on-demand runtime workspaces writing namespace mirrors at all: each provision reaches a discarded filesystem under a fresh address, so the entry can never be read back and only spends a capped slot that a local or SSH project needs. The repo-id row still records the name for the live session. * fix(worktrees): re-insert migrated namespaces so the cap cannot undo a migration Exempting keys from the trim protected them for that one call and no other. A merged destination keeps its original insertion slot, so it sat at the front of the eviction queue and the next unrelated retirement write dropped it — losing both the migrated name and the name the destination already held, on a host that had just been re-added. Re-insert what the migration writes instead, the same discipline the ordinary writer already follows, so insertion order reflects use. That also removes the exemption, which could otherwise leave the map stuck at twice the cap until one later write evicted the whole excess at once. Corrects the runtime-gate comment as well: the mirror is unreadable after the next provision, not immediately, so a remove/re-add inside one provision is a real if narrow loss. * fix(worktrees): refresh a retained namespace source even when its merge adds nothing Replacing the trim exemption with re-insertion narrowed the protection: the exemption covered every retained source, the re-insertion only covered sources whose merge actually wrote. A copy whose destination already held the same names was then neither re-inserted nor exempt, so the migration's own trim evicted the shared source bucket ahead of hundreds of untouched ones — losing the tombstones of a live sibling target still on that endpoint, which is what copying exists to prevent. A move's destination gets the same treatment: deleting the source makes it the only remaining copy, so it has been used. Both are order-only and deliberately do not set the changed flag, keeping an import that moved nothing from scheduling a save.
212 lines
9.0 KiB
TypeScript
212 lines
9.0 KiB
TypeScript
import { getRepoExecutionHostId, parseExecutionHostId } from '../shared/execution-host'
|
|
import {
|
|
mergeRetiredNameRegistries,
|
|
type RetiredNameRegistry
|
|
} from '../shared/worktree/retired-name-registry'
|
|
import type { Repo } from '../shared/repo-types'
|
|
import { worktreePathComparisonKey } from './ipc/worktree-path-comparison'
|
|
import { sshEndpointKey, type SshIdentityFields } from './ssh/ssh-target-identity'
|
|
|
|
/**
|
|
* Keys for `retiredWorktreeNamesByNamespace`, the copy of the retirement registry that survives a
|
|
* project being removed and re-added.
|
|
*
|
|
* Primary storage stays on `repo.id` (see worktree-name-retirement.ts): the namespace is derived
|
|
* from settings that a user can toggle at any time, so keying storage on it would orphan every
|
|
* affected repo's retirements on a settings change. This map is the second, path-derived copy —
|
|
* a settings toggle changes this key but `repo.id` still holds the tombstone, and a remove/re-add
|
|
* loses the `repo.id` row but this key still matches.
|
|
*/
|
|
|
|
type RepoHostFields = Pick<Repo, 'connectionId' | 'executionHostId'>
|
|
|
|
export type SshTargetLookup = (targetId: string) => SshIdentityFields | undefined
|
|
|
|
/** No target row means no way to tell endpoints apart. The retirement contract prefers
|
|
* over-retiring (one name out of a 552-name pool) to reissuing a path whose agent history is
|
|
* still there, so unresolvable SSH repos share a bucket — still split by workspace path. */
|
|
export const UNKNOWN_SSH_HOST_IDENTITY = 'ssh:?'
|
|
|
|
/** Which machine and account a repo creates workspaces on. SSH resolves to the endpoint tuple
|
|
* rather than the target row id, because a removed and re-added host mints a fresh id while
|
|
* reaching the same filesystem. */
|
|
export function retirementHostIdentity(repo: RepoHostFields, lookup?: SshTargetLookup): string {
|
|
const hostId = getRepoExecutionHostId(repo)
|
|
const parsed = parseExecutionHostId(hostId)
|
|
if (parsed?.kind !== 'ssh') {
|
|
return hostId
|
|
}
|
|
const target = lookup?.(parsed.targetId)
|
|
return target ? sshHostIdentity(target) : UNKNOWN_SSH_HOST_IDENTITY
|
|
}
|
|
|
|
export function sshHostIdentity(target: SshIdentityFields): string {
|
|
return `ssh:${sshEndpointKey(target)}`
|
|
}
|
|
|
|
export function retirementNamespaceKey(hostIdentity: string, probePath: string): string {
|
|
return `${hostIdentity}:${worktreePathComparisonKey(probePath)}`
|
|
}
|
|
|
|
/** Rewrites a key's host identity, keeping its workspace-path half. Returns null when the key is
|
|
* not under `fromIdentity` or the swap is a no-op. */
|
|
function swapRetirementNamespaceHost(
|
|
namespaceKey: string,
|
|
fromIdentity: string,
|
|
toIdentity: string
|
|
): string | null {
|
|
const prefix = `${fromIdentity}:`
|
|
return fromIdentity !== toIdentity && namespaceKey.startsWith(prefix)
|
|
? `${toIdentity}:${namespaceKey.slice(prefix.length)}`
|
|
: null
|
|
}
|
|
|
|
/** The canonical key plus its pre-identity twin, whose host half was the SSH target row id. Reads
|
|
* cover both so an upgrade keeps the tombstones it already wrote; writes only use the first. */
|
|
export function retirementNamespaceKeysToRead(
|
|
repo: RepoHostFields,
|
|
namespaceKey: string,
|
|
lookup?: SshTargetLookup
|
|
): string[] {
|
|
const legacyKey = swapRetirementNamespaceHost(
|
|
namespaceKey,
|
|
retirementHostIdentity(repo, lookup),
|
|
getRepoExecutionHostId(repo)
|
|
)
|
|
return legacyKey ? [namespaceKey, legacyKey] : [namespaceKey]
|
|
}
|
|
|
|
/** The map deliberately outlives the repos that wrote it — that is what makes a re-add recover its
|
|
* tombstones — so nothing prunes it per repo. It is capped instead. */
|
|
export const MAX_RETIREMENT_NAMESPACES = 256
|
|
|
|
export function recordRetirementNamespaceRegistry(
|
|
namespaces: Record<string, RetiredNameRegistry>,
|
|
namespaceKey: string,
|
|
registry: RetiredNameRegistry
|
|
): void {
|
|
// Re-inserting moves the key to the end: JS keeps non-numeric string keys in insertion order and
|
|
// JSON round-trips it, so the object's own order is the LRU list — no timestamp to persist.
|
|
delete namespaces[namespaceKey]
|
|
namespaces[namespaceKey] = registry
|
|
trimRetirementNamespaces(namespaces)
|
|
}
|
|
|
|
/** Evicts oldest-first back to the cap. Every writer that can grow the map must end here — a copy
|
|
* during migration adds a key without removing one, so the map would otherwise drift over the cap
|
|
* and the next ordinary write would evict the whole excess at once.
|
|
*
|
|
* Callers are responsible for re-inserting anything they touched first, so that insertion order
|
|
* actually reflects use. Exempting keys here instead would protect them for this one call and no
|
|
* other, leaving a just-written bucket sitting at the front of the eviction queue. */
|
|
function trimRetirementNamespaces(namespaces: Record<string, RetiredNameRegistry>): void {
|
|
const keys = Object.keys(namespaces)
|
|
// Why the clamp: a negative `end` makes slice count back from the tail, so an under-full map
|
|
// would evict almost everything it holds instead of nothing.
|
|
const overflow = Math.max(0, keys.length - MAX_RETIREMENT_NAMESPACES)
|
|
for (const stale of keys.slice(0, overflow)) {
|
|
delete namespaces[stale]
|
|
}
|
|
}
|
|
|
|
export type RetirementNamespaceHostMigration = {
|
|
/** Identities a single target row owned outright — its `ssh:<targetId>` host id. Reassignment
|
|
* leaves nothing pointing at them, so their namespaces move. */
|
|
moveFrom?: readonly string[]
|
|
/** Endpoint identities. These are deliberately *not* owned by the row that rotated: nothing stops
|
|
* a second target resolving to the same host|port|username, so the source bucket is kept.
|
|
* Copying costs one duplicated entry; moving would strip a live target's tombstones and reissue
|
|
* a path whose agent history is still on disk. */
|
|
copyFrom?: readonly string[]
|
|
to: string
|
|
}
|
|
|
|
/** Re-keys every namespace an SSH target used to reach onto its current identity, so a rotated
|
|
* target id does not strand the names it already spent. */
|
|
export function migrateRetirementNamespaceHostIdentity(
|
|
namespaces: Record<string, RetiredNameRegistry> | undefined,
|
|
migration: RetirementNamespaceHostMigration
|
|
): boolean {
|
|
if (!namespaces) {
|
|
return false
|
|
}
|
|
const visited = new Set<string>()
|
|
let changed = false
|
|
for (const group of [
|
|
{ identities: migration.moveFrom ?? [], retainSource: false },
|
|
{ identities: migration.copyFrom ?? [], retainSource: true }
|
|
]) {
|
|
for (const oldIdentity of group.identities) {
|
|
if (!oldIdentity || oldIdentity === migration.to || visited.has(oldIdentity)) {
|
|
continue
|
|
}
|
|
visited.add(oldIdentity)
|
|
if (rekeyRetirementNamespaceHost(namespaces, oldIdentity, migration.to, group.retainSource)) {
|
|
changed = true
|
|
}
|
|
}
|
|
}
|
|
if (changed) {
|
|
// A retained source bucket grows the map, so migration has to respect the cap too.
|
|
trimRetirementNamespaces(namespaces)
|
|
}
|
|
return changed
|
|
}
|
|
|
|
function retiredNameRegistriesEqual(a: RetiredNameRegistry, b: RetiredNameRegistry): boolean {
|
|
if (a.exhaustedTiers !== b.exhaustedTiers || a.names.length !== b.names.length) {
|
|
return false
|
|
}
|
|
const names = new Set(a.names)
|
|
return b.names.every((name) => names.has(name))
|
|
}
|
|
|
|
function rekeyRetirementNamespaceHost(
|
|
namespaces: Record<string, RetiredNameRegistry>,
|
|
fromIdentity: string,
|
|
toIdentity: string,
|
|
retainSource: boolean
|
|
): boolean {
|
|
const prefix = `${fromIdentity}:`
|
|
let changed = false
|
|
for (const key of Object.keys(namespaces)) {
|
|
if (!key.startsWith(prefix)) {
|
|
continue
|
|
}
|
|
const registry = namespaces[key]
|
|
const nextKey = `${toIdentity}:${key.slice(prefix.length)}`
|
|
const existing = namespaces[nextKey]
|
|
// Why compare: a copy repeats on every import, and reporting a change the destination already
|
|
// covered would schedule a save each time. Compared by membership rather than by size, so a
|
|
// destination that was stored uncompacted cannot trade a folded name for a new one and read
|
|
// as unchanged.
|
|
const merged = existing ? mergeRetiredNameRegistries(existing, registry) : registry
|
|
const wrote = !existing || !retiredNameRegistriesEqual(merged, existing)
|
|
if (!retainSource) {
|
|
delete namespaces[key]
|
|
changed = true
|
|
} else {
|
|
// Unconditional, even when the merge added nothing: a retained source is the bucket a live
|
|
// sibling target still reads, so it counts as used. Gating this on a write would let the trim
|
|
// take it ahead of buckets nothing has touched in months. Order only — never `changed`, or an
|
|
// import that moved nothing would schedule a save.
|
|
delete namespaces[key]
|
|
namespaces[key] = registry
|
|
}
|
|
if (wrote) {
|
|
// Re-insert rather than assign: assigning to an existing key leaves it in its original slot,
|
|
// so a destination that was just written would sit at the front of the eviction queue and the
|
|
// next ordinary write would drop it — taking the migrated name and its own with it.
|
|
delete namespaces[nextKey]
|
|
namespaces[nextKey] = merged
|
|
changed = true
|
|
} else if (existing) {
|
|
// A move just made this destination the only copy, and a no-op copy is still a use. Same
|
|
// reasoning as above, and likewise order only.
|
|
delete namespaces[nextKey]
|
|
namespaces[nextKey] = existing
|
|
}
|
|
}
|
|
return changed
|
|
}
|