Files
orca/cloud/dev
Jinwoo-H e39f4ebad5 fix(relay): compare hinted regions against placed ones, not a fixed share
Review found the skew alert inverted at both ends. A fixed 40% bar on the
asia-east2 share of region hints was silent through the exact broken state it
was written for, and would page forever once the desktop probe is fixed and
the genuine APAC share rises past it. An absolute share cannot separate those
because it has no reference point.

The hint share now has one: the share of assignments the director actually
placed in that region during the same hour. Measured over twelve hours on
2026-09-07, while the probe was still mis-picking, asia-east2 was 33.8% of
33,800 hinted requests and 7.9% of 45,364 assignments. That is a 4.27x
divergence and a 25.9-point gap, so the alert fires above 2x and 15 points,
inside the broken state and outside a healthy one. Both bars must hold: the
ratio alone blows up on tiny placement counts, the gap alone misses a
proportionally large skew at low volume. The reviewer proposed either bar
alone; requiring both keeps each one meaningful and still clears today's
numbers with room.

`unhinted` requests leave the denominator. They were 27% of all requests, so
a client that always sends a hint would move the number from 21.9% to 35.0%
with no behaviour change at all.

The comparison needs per-region placement counters, so `selectedRegionsDelta`
gets log-based metrics alongside the requested ones. Rather than extract four
hyphenated map keys through quoted field paths, which nothing in the project
does and which cannot be checked without applying, the relay now also
publishes flat `requestedRegion<Region>Delta` and `selectedRegion<Region>Delta`
fields next to the untouched maps. They are emitted as zeros in every
interval, so no series can drop out of the alert's inner join in an hour with
no asia placements, which is exactly the hour the skew is worst. Additive
only: metricVersion is unchanged, the maps still carry anything outside the
catalog, and the emitter's leak guard still passes.

Two corrections to what the previous commit claimed. None of these metrics
exist in the project yet, so the code, the doc and this message now say what
was actually checked against production: the query shapes, run over existing
metrics of the same kind. And the control-RTT policy records that EU desktops
on us-central1 sit at 100-130 ms, so a European-heavy cell can approach the
150 ms bar while correctly homed.

The skew alert will stay lit after a client fix until the backlog is rehomed.
Sticky assignment never re-consults the hint, so a desktop already on an asia
cell keeps landing there whatever it now asks for. The policy description and
the doc both say so, so nobody reads a slow clear as a failed fix.
2026-09-07 04:41:09 -04:00
..