Files
orca/src/shared/pty-consumer-session-hello.ts
Jinwoo Hong eea0bb64db fix(ssh): make PTY owner admission explicit and non-destructive (#12673)
An owner-capable `pty.openClient` had two failure modes that presented as something else.

If the relay still held an owner record but the request carried no matching resume proof, admission fell through to a SUBSCRIBER grant — a success-shaped response the client cannot use, which it then rejected as "did not grant an authenticated PTY session owner". And if the relay had forgotten the record the client named, admission threw a stale-recovery error, which the client answered by deleting its own recovery row — `clientInstanceId` included — and reopening. Two round trips, and the identity that lets it resume that target at all went with the deletion.

Now every owner grant carries a required `resumed` flag, a forgotten record mints a fresh claim in one round trip, a held claim returns one of three coded refusals, duplicate opens on one connection are rejected even when identical, and an attached-holder refusal becomes a typed error routed through the terminal-relay-error callback instead of feeding redeploy backoff a link that is working fine.

Independent review caught two regressions in the first attempt, both now fixed and both with tests that fail without them:

**A backpressure teardown could take a live owner's session.** The safety argument was that a record only becomes `disconnected` from an observed peer close — but two of the six paths there are capacity paths, where the relay destroys the client's socket itself because its lane queue filled. That is the signature of a client that is ALIVE but not draining fast enough. Demonstrated: the real owner is torn down for backpressure, a rival is granted ownership 270ms into a nominal 30s grace, and the owner's later reconnect with a valid resume proof is refused permanently, backoff cleared, no retry. Closes now carry a cause (`peer-closed` | `local`, defaulting to `local`, which only ever widens a grace), and the floor applies only to closes the transport actually observed on the peer's side. Capacity teardowns, decode faults and sink failures keep the default.

**A client's own zombie connection blocked it permanently.** Only `SshRelaySession` ever requests owner, and every endpoint-credential client shares one principal — so in a normal single-app deployment an `active` incumbent refusing you is almost always your own half-open connection the relay never saw close. That was refused as terminal, where main recovered on bounded backoff once keepalive noticed. The refusal already held both client identities; a match is now a distinct transient refusal that falls through to relay-lost backoff, restoring that recovery. A genuinely different client is still blocked.

Also: each retry deadline now starts when its own phase begins, instead of both being computed at entry where a slow first phase could leave the second with zero attempts.

Fixes STA-3365.
2026-08-05 02:56:07 -07:00

37 lines
1.5 KiB
TypeScript

import type { PtyConsumerSessionHello } from './pty-consumer-session-contract'
export const MAX_CAPABILITY_VERSIONS = 8
export function assertNonEmptyString(value: unknown, name: string): asserts value is string {
if (typeof value !== 'string' || value.length === 0 || value.length > 512) {
throw new Error(`${name} must be a non-empty string of at most 512 characters`)
}
}
export function validateHello(hello: PtyConsumerSessionHello): void {
assertNonEmptyString(hello.clientInstanceId, 'clientInstanceId')
if (hello.requestedRole !== 'session-owner' && hello.requestedRole !== 'subscriber') {
throw new Error('requestedRole must be session-owner or subscriber')
}
if (hello.resume) {
if (!Number.isSafeInteger(hello.resume.ownerGeneration) || hello.resume.ownerGeneration <= 0) {
throw new Error('resume.ownerGeneration must be a positive safe integer')
}
assertNonEmptyString(hello.resume.ownerLease, 'resume.ownerLease')
}
const flow = hello.capabilities?.outputFlowControl
if (!flow) {
return
}
if (
!Array.isArray(flow.versions) ||
flow.versions.length > MAX_CAPABILITY_VERSIONS ||
flow.versions.some((version) => !Number.isSafeInteger(version) || version <= 0)
) {
throw new Error('outputFlowControl.versions must contain positive safe integers')
}
if (!Number.isSafeInteger(flow.requestedWindowSu) || flow.requestedWindowSu <= 0) {
throw new Error('outputFlowControl.requestedWindowSu must be a positive safe integer')
}
}