mirror of
https://github.com/stablyai/orca.git
synced 2026-10-09 00:02:39 +00:00
* fix(relay): refuse host hellos fast when the database pool is saturated; renewals jump the queue A reconnect herd (2026-10-06 18:34Z, 1,046 asia-east2 hosts dropped by the load balancer at once) queued hundreds of host hellos behind each cell's 16-connection pool. They timed out after 2 s, the desktops redialled into the same queue, and control-lease renewals starved behind them. - Host control upgrades get an immediate 503 (Retry-After: 2) once 32 callers wait for a pooled connection. Shipped desktops already treat that as a connect error and back off with jittered exponential retry. - General work now queues in PostgresPoolPressure instead of pg-pool, so renewal statements (priority lane) take the next released connection. - A renewal batch that never got a connection no longer fans out into one statement per host. - hostHellosShedDelta counts refusals in the runtime metrics. * fix(relay): shed host hellos on the pool's oldest wait, not its queue length asia-east2 queues of 50-196 waiters are routine and keep moving, so a 32-waiter limit would refuse hellos that were going to succeed. Shed only while the oldest pool waiter has waited 1.5 s of the 2 s acquire timeout, when a new hello would time out anyway; a refused desktop gets the same connect error and backoff it gets today, 2 s sooner. Rebinds over a live control (lease rotation) are never refused. * fix(relay): shed hellos only when a full pool queue has aged 1 s Per review: require both an oldest waiter of at least 1 s and a pool's worth of waiters, so one slow waiter cannot trip it, and state the real baseline (p99 of 2 waiters on 10-07).