Files
orca/cloud/apps/relay
Jinwoo Hong 3062b9d8d0 fix(relay): refuse host hellos fast when the database pool is timing them out; renewals jump the queue (#26362)
* fix(relay): refuse host hellos fast when the database pool is saturated; renewals jump the queue

A reconnect herd (2026-10-06 18:34Z, 1,046 asia-east2 hosts dropped by the
load balancer at once) queued hundreds of host hellos behind each cell's
16-connection pool. They timed out after 2 s, the desktops redialled into
the same queue, and control-lease renewals starved behind them.

- Host control upgrades get an immediate 503 (Retry-After: 2) once 32
  callers wait for a pooled connection. Shipped desktops already treat
  that as a connect error and back off with jittered exponential retry.
- General work now queues in PostgresPoolPressure instead of pg-pool, so
  renewal statements (priority lane) take the next released connection.
- A renewal batch that never got a connection no longer fans out into
  one statement per host.
- hostHellosShedDelta counts refusals in the runtime metrics.

* fix(relay): shed host hellos on the pool's oldest wait, not its queue length

asia-east2 queues of 50-196 waiters are routine and keep moving, so a
32-waiter limit would refuse hellos that were going to succeed. Shed only
while the oldest pool waiter has waited 1.5 s of the 2 s acquire timeout,
when a new hello would time out anyway; a refused desktop gets the same
connect error and backoff it gets today, 2 s sooner. Rebinds over a live
control (lease rotation) are never refused.

* fix(relay): shed hellos only when a full pool queue has aged 1 s

Per review: require both an oldest waiter of at least 1 s and a pool's worth
of waiters, so one slow waiter cannot trip it, and state the real baseline
(p99 of 2 waiters on 10-07).
2026-10-07 21:18:26 -04:00
..