Design Counsel — Terminal Session Ownership

Three review rounds across two models, an adversarial fact-check, and a close-out · the design that emerged is not the one I proposed

1
Live bug found
2 → 1
Durable records
3
Review rounds
78 / 12
Facts right / wrong
4
Unresolved conflicts
~900–1250
Lines deleted

The headline

The counsel found a live duplicate-agent bug reachable from a shipped gesture, and then something worse: respawn is reachable down at least five paths, and the proof gate I shipped sits on only one of them — not the primary one. Detaching a pane into a new tab makes a healthy remote shell report as dead, and the client starts a second one with agent resume while the original keeps running.

1 · How the counsel ran

Process and convergence⤢ zoom

Round 1

Grok · 10 findings

Opus · 8 findings

GPT seat · empty

shared brief
verified facts vs unverified leads

synthesis
both models independently say:
one record, not two

Round 2 · Grok 6 · Opus 7
→ key by leafId, not paneKey

Round 3 · Grok 5 · Opus 5
→ delete attach identity entirely

clean-check + fact-check

close-out
C1–C5

Round 1

Grok · 10 findings

Opus · 8 findings

GPT seat · empty

shared brief
verified facts vs unverified leads

synthesis
both models independently say:
one record, not two

Round 2 · Grok 6 · Opus 7
→ key by leafId, not paneKey

Round 3 · Grok 5 · Opus 5
→ delete attach identity entirely

clean-check + fact-check

close-out
C1–C5

The GPT seat returned nothing in all three rounds and was dropped. Grok ran through its CLI; the wrapper was instructed to report a failure rather than author findings, so its participation is real rather than assumed. Findings fell 10 → 6 → 5 and 8 → 7 → 5 across rounds — converging, not exhausted.

2 · The live bug

Four independently reasonable decisions compose into a shell-killer.

Detach a pane → duplicate agent⤢ zoom

user detaches a pane
into a new tab
shipped gesture

tabId changes;
relay still holds the old one
frozen at spawn

reattach → relay compares
and throws
PTY "x" not found (identity mismatch)

isSshPtyNotFoundError
regex /PTY ".+" not found/
matches

re-emitted as
SSH_SESSION_EXPIRED

isProvenSshSessionGoneError
returns true

clear binding →
startFreshColdRestoreAgentResume

second shell + SECOND AGENT RESUME
original shell still running

user detaches a pane
into a new tab
shipped gesture

tabId changes;
relay still holds the old one
frozen at spawn

reattach → relay compares
and throws
PTY "x" not found (identity mismatch)

isSshPtyNotFoundError
regex /PTY ".+" not found/
matches

re-emitted as
SSH_SESSION_EXPIRED

isProvenSshSessionGoneError
returns true

clear binding →
startFreshColdRestoreAgentResume

second shell + SECOND AGENT RESUME
original shell still running

The error string says the PTY was not found. It was found — that is how its identity got compared. The word “not found” is doing double duty for “absent” and “present but mismatched”, and a regex cannot tell them apart.

// my classifier — the last link
export function isProvenSshSessionGoneError(error: unknown): boolean {
  if (message.includes(SSH_SOURCE_RESTORE_REQUIRED_ERROR)) return false
  return message.includes(SSH_SESSION_EXPIRED_ERROR) || /PTY ".+" not found/i.test(message)
}
// two functions away, unused by me:
export function isSshPtyIdentityMismatchError(error: unknown): boolean

Why I missed it

I wrote this file to enforce “unknown is not dead,” and its doc-comment says “True only when the failure proves the session no longer exists.” I treated a message shape as proof. The discriminator I needed already existed in the same file. A test asserting “not-found ⇒ gone” passes happily, because the wrongness is in what the string means, not in what the code does with it.

The gate is on the wrong path

I shipped “respawn requires positive proof” and pinned it with a test. Grok traced every respawn path and found the proof check is bypassed on the main one, because the error never reaches a classifier — it is converted into a boolean first:

// pty-transport.ts — catches the error and RESOLVES
return { id: options.sessionId, sessionExpired: true }

// pty-connection.ts — respawns on the flag alone, no proof check
if (connectResult?.sessionExpired) { /* clear binding */
  startFreshColdRestoreAgentResume(...) }
Respawn paths vs the proof gate⤢ zoom

resolved as a flag

bare token

empty result

main process

thrown

reattach fails

how does the
failure travel?

sessionExpired: true
→ respawn
no proof check

isSshSessionExpiredError
→ respawn

no ptyId returned
→ respawn

lease marked expired
→ recoverTerminalPane

isProvenSshSessionGoneError
the gate I shipped

resolved as a flag

bare token

empty result

main process

thrown

reattach fails

how does the
failure travel?

sessionExpired: true
→ respawn
no proof check

isSshSessionExpiredError
→ respawn

no ptyId returned
→ respawn

lease marked expired
→ recoverTerminalPane

isProvenSshSessionGoneError
the gate I shipped

What this means for the shipped work

Two of my six changes are now known to be narrower than claimed. The keystroke fence is inert on the reattach path, and the respawn gate guards a minority of respawn routes. Both passed their tests because each test asserted the guard behaves correctly when reached — neither could observe that the guarded route is not the route production takes.

3 · What the counsel changed about my design

My proposalCounsel’s designWhy theirs is better
Two records: ownership + orphan inventoryOne record. Orphans become a connect-time projection with no durable rowsA durable orphan row goes stale across exactly the disconnect during which the truth changed — and an orphan is only actionable while connected. Also deletes the two-record transaction my split needed.
Key by paneKey = tabId:leafIdKey by leafId; tab derived from the live layout, never storedA stored tabId is either frozen (names a tab the leaf left) or updated (desyncs from the relay’s copy). Both are live failures. Keying by the leaf deletes the transfer problem instead of managing it.
Keep attach-time pane identityDelete it — compare returned incarnation insteadThe apparatus is over-strict and under-strict at once: it rejects a live renamed pane, and accepts a different shell after a relay restart recycles the id. It fails at its own stated job.
“Unknown is not dead” as a principleSame rule, made structural: fate must be reported by the host, never inferred from timing, absence, or error-string shapeEvery defect found across three rounds is one violation of that clause. Mine was a slogan; theirs is a rule with an enforcement point.

Both models proposed the single-record collapse independently, which is the strongest signal available in a counsel and why it outweighed my version.

4 · The rule everything follows from

Ownership vs fate⤢ zoom

FATE — is the shell alive

non-durable

host-local

must be REPORTED by the host
never INFERRED by the client

OWNERSHIP — which pane owns which shell

durable

client-local

read synchronously
from disk while offline

not from timing
not from an absent entry
not from an error string

FATE — is the shell alive

non-durable

host-local

must be REPORTED by the host
never INFERRED by the client

OWNERSHIP — which pane owns which shell

durable

client-local

read synchronously
from disk while offline

not from timing
not from an absent entry
not from an error string

This is what makes the offline case work without an authority layer. Ownership is answerable with the network down, because it never depended on the network. Fate is never answered by the client at all.

5 · The record

/** Main-owned. TOP-LEVEL in PersistedState — never inside WorkspaceSessionState. */
type PaneShellOwnership = {
  leafId: string          // PRIMARY KEY — layout leaf UUID
  hostId: ExecutionHostId // FIELD: 'local' | 'ssh:<target>' | 'runtime:<id>'
  worktreeId: string
  ptyId: string           // relay-native form for SSH — one namespace, always
  incarnationId: string   // REQUIRED
  relayInstanceId?: string
  state: 'attached' | 'detached' | 'terminated'
  createdAt / updatedAt / lastAttachedAt? / lastDetachedAt?
}
// Absent on purpose: no tabId, no paneKey, no hostEpoch, no second record.

Uniqueness is enforced on (hostId, ptyId, incarnationId), not ptyId alone — because relay ids are minted pty-1, pty-2… per relay process, so pty-1 recurs on every host after every relay restart. Round 1 had rejected a collision here as impossible; round 3 showed it is guaranteed, and the rejection was withdrawn.

6 · Orphans are a projection, not an authority

Why no durable orphan rows⤢ zoom

rejected

on connect

ask the host what it has

subtract what our records claim
incarnation-scoped

remainder = orphans
computed, never stored

feeds the existing cleanup affordance
never auto-killed

a durable orphan row would
go stale across the very disconnect
during which the truth changed

rejected

on connect

ask the host what it has

subtract what our records claim
incarnation-scoped

remainder = orphans
computed, never stored

feeds the existing cleanup affordance
never auto-killed

a durable orphan row would
go stale across the very disconnect
during which the truth changed

7 · What gets deleted

DeletedWhere
Attach-identity apparatusrelay comparison + throw, the error constants, the expected-identity derivation on both reattach paths, the renderer branch that treats a mismatch as proof
Lease↔binding reconciliationsupersession, ranking, deferral, rollback, dual matchers — the ~209 lines I added, plus the ~374 that preceded them
The restorable forkthe two contradictory predicates where absence of a lease meant opposite things depending on a sibling leaf
The 30-second grant arithmeticreplaced by an explicit record state
Never built at allhostEpoch and its clock arithmetic; a five-row evidence matrix; a resumed flag — all proposed in earlier rounds and killed before implementation

Net lines, honestly

Removed ≈900–1,250. Added ≈450–700. Net ≈ −300 to −700 — but only after a one-release legacy projection is dropped; during that release it is ≈ −100 to −450. Tests go up ≈900–1,500 for 25 oracles, several needing SSH + detach + relay-restart harnesses that do not exist. The synthesizer stated plainly it could not narrow the range further without doing the call-site rewiring.

8 · Verification

Models assert plausible file contents confidently, so every claim about existing code was checked.

78
Confirmed
12
Wrong
1
Unverifiable

Eleven of the twelve errors were citation-level and changed no conclusion. One was a real premise defect: the claim that abandoning a wrongly-attached shell is free. It is not — attach commits activation, clears pending output and may replay before it returns the incarnation the client would use to detect the mistake, and there is no detach RPC to undo it. That correction is folded in as C4.

I independently verified the duplicate-agent chain end to end, since it implicates my own code.

9 · Where the models disagreed

ConflictStatus
Do shells survive a hard relay death on Windows/WSL?Open — adjudicated toward “dead” on POSIX evidence; the Windows case is not established
Does the 30-second recovery grant execute in production at all?Open — two reviewers read its comparison as unreachable for a real SSH pane; the covering test seeds a shape production cannot produce
Should relayInstanceId exist?Open — one would delete it and let incarnation carry the rule; the other keeps it because a re-minted id proves only that the relay restarted
Fix or delete the attach identity?Resolved — deletion, and the synthesizer went further than either reviewer

10 · Readiness

A final pass re-checked whether the close-out actually closed its own blockers. Verdict: STILL NOT READY.

BlockerClose-out
The spawn reattach path still wraps any not-found as expired, so client-side auto-respawn survivesNOT CLOSED — C1 hardens the classifier, but the primary path never reaches a classifier. Every auto-respawn consumer must require proof, not just the one that catches a throw.
Migrated records have no incarnation, but the design requires one and throws on mismatchCLOSED, conditionally — C2/C3 work only if an absent recorded incarnation means learn and stamp, never mismatch. That rule must be written into both the algorithm and the type.

11 · Decisions for you

#Decision
1The live bug. It is independent of the redesign and fixable in isolation. Fix it now as its own change, or fold it into the track?
2Visible detached panes. The design replaces an invisible 30-second auto-recovery with an explicit state and two actions (reattach / start a new shell). That is a deliberate product change — a user-initiated relay reset would now show detached panes rather than silently recreating shells.
3Release policy. A leaf removed while its shell lives releases the claim — the shell becomes orphan-eligible and adoptable, and is never auto-killed. Confirm you want liveness preserved over tidiness.
4Sequencing. Two preconditions block everything: one partition per (target, pane), and one bind producer. The second is also what would make the shipped fence non-inert.
Companion to report.html (evidence for the shipped work) and finalized-design.html (the two-plane decision this supersedes at the control plane). Click any diagram to zoom · scroll to scale · drag to pan.
×
scroll to zoom · drag to pan · Esc to close