mirror of
https://github.com/stablyai/orca.git
synced 2026-10-09 08:02:35 +00:00
* feat(relay): alert on cell disconnect bursts and database pool herds Two alert policies that separate a load-balancer mass disconnect from a database stall, after the 2026-10-06 18:34Z asia-east2 drop: - Disconnect burst: over 120 code-1006 desktop control closes on one cell in a minute, counted from the cell's own close line. The LB's internal_error log entries carry the stream start as their timestamp (median 2.4 h before the drop), so an LB-log rate never shows the pulse. - Pool herd: a cell runtime sample with more than 50 requests queued for a PostgreSQL connection (interval maximum, bar in the filter). Also publishes orca_relay_host_hellos_shed for the refusal counter from #26362, and pins both filters to the relay's text in a contract test. * fix(relay): raise the cell pool-herd alert bar to 200 waiters asia-east2 cells reach 184-196 waiters outside any herd, so >50 would have fired in 84 episodes since 2026-10-03. >200 fires in 10, covering every known herd. Drops the false 'healthy peak is 2 waiters' baseline from the policy text. * docs(relay): describe pool-herd baseline as stalls, not routine load A healthy cell peaks at 2 waiters (p99); the 50-196 readings are single-cell stalls that already time out. The 200 bar is unchanged.