Files
orca/cloud/dev
Jinwoo Hong 8cf6228655 feat(relay): alert on cell disconnect bursts and database pool herds (#26363)
* feat(relay): alert on cell disconnect bursts and database pool herds

Two alert policies that separate a load-balancer mass disconnect from a
database stall, after the 2026-10-06 18:34Z asia-east2 drop:

- Disconnect burst: over 120 code-1006 desktop control closes on one cell
  in a minute, counted from the cell's own close line. The LB's
  internal_error log entries carry the stream start as their timestamp
  (median 2.4 h before the drop), so an LB-log rate never shows the pulse.
- Pool herd: a cell runtime sample with more than 50 requests queued for a
  PostgreSQL connection (interval maximum, bar in the filter).

Also publishes orca_relay_host_hellos_shed for the refusal counter from
#26362, and pins both filters to the relay's text in a contract test.

* fix(relay): raise the cell pool-herd alert bar to 200 waiters

asia-east2 cells reach 184-196 waiters outside any herd, so >50 would have
fired in 84 episodes since 2026-10-03. >200 fires in 10, covering every known
herd. Drops the false 'healthy peak is 2 waiters' baseline from the policy text.

* docs(relay): describe pool-herd baseline as stalls, not routine load

A healthy cell peaks at 2 waiters (p99); the 50-196 readings are single-cell
stalls that already time out. The 200 bar is unchanged.
2026-10-07 21:03:49 -04:00
..